BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055402Z
LOCATION:231-232
DTSTART;TZID=America/Chicago:20211115T110000
DTEND;TZID=America/Chicago:20211115T113000
UID:submissions.supercomputing.org_SC21_sess343_ws_h2rc110@linklings.com
SUMMARY:Near-Data FPGA-Accelerated Processing of Collective and Inference 
 Operations in Disaggregated Memory Systems
DESCRIPTION:Workshop\n\nNear-Data FPGA-Accelerated Processing of Collectiv
 e and Inference Operations in Disaggregated Memory Systems\n\nHeinz, Koch\
 n\nWith growing data set sizes, many scientific and data center HPC worklo
 ads observe an increasing scaling imbalance, e.g., between compute and mem
 ory capacities. As a solution, disaggregated system architectures employ s
 patial distribution of the different resources. They aim for independent s
 caling of the different resource kinds (e.g., compute, non-volatile storag
 e, memory), and use fast communication fabrics for their interconnection.\
 n\nHowever, for some bulk operations, such as reductions and collections, 
 it is still beneficial to perform them close to the memories, avoiding the
  need to move large volumes of data over the fabric. \n\nThis work realize
 s a disaggregated system capable of performing such near-data processing (
 NDP) operations by extending the distributed memory controllers with hardw
 are-accelerated compute capabilities. The actual computations execute on F
 PGAs and can be abstractly described using C/C++ as compilable by high-lev
 el hardware synthesis (HLS) tools.\n\nWe have aimed for high usability of 
 our technology also by HPC experts unfamiliar with hardware design. An aut
 omated toolflow encapsulates the creation and deployment of the actual acc
 elerators in the disaggregated system. The NDP operations execute distribu
 ted across all memory nodes, and are easily accessed using a simple MPI-ba
 sed programming interface that requires only minimal effort to use in exis
 ting applications.\n\nOur solution is demonstrated using a prototype disag
 gregated system based on the low-latency EXTOLL fabric for communication. 
 We evaluate both conventional reductions/collectives as well as complete m
 achine-learning inference tasks.\n\nTag: Accelerator-based Architectures, 
 Applications, Architectures, Emerging Technologies, Heterogeneous Systems,
  Memory Systems, Networks\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
