BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055402Z
LOCATION:231-232
DTSTART;TZID=America/Chicago:20211115T140000
DTEND;TZID=America/Chicago:20211115T143000
UID:submissions.supercomputing.org_SC21_sess343_ws_h2rc113@linklings.com
SUMMARY:Optimized Implementation of the HPCG Benchmark on Reconfigurable H
 ardware
DESCRIPTION:Workshop\n\nOptimized Implementation of the HPCG Benchmark on 
 Reconfigurable Hardware\n\nZeni\n\nThe HPCG benchmark represents a modern 
 complement to the HPL benchmark in the performance evaluation of HPC syste
 ms, as it has been recognized as a more representative benchmark to reflec
 t real-world applications. While typical workloads become more and more ch
 allenging, the semiconductor industry is battling with performance scaling
  and power efficiency on next-generation technology nodes. As a result, th
 e industry is turning towards more customized compute architectures to hel
 p meet the latest performance requirements. In this paper, we present the 
 details of the first FPGA-based implementation of HPCG that takes advantag
 e of such customized compute architectures. Our results show that our high
 -performance multi-FPGA implementation, using 1 and 4 Xilinx Alveo U280 ac
 hieves up to 108.3 GFlops and 346.5 GFlops respectively, representing spee
 d-ups of   104.1×  and   333.2×  over software running on a server with an
  Intel Xeon processor with no loss of accuracy. We also demonstrate that t
 he FPGA-based solution achieves comparable performance with respect to mod
 ern GPUs and an up to   2.7×  improvement in terms of power efficiency com
 pared to an NVIDIA Tesla V100. Finally, a theoretical evaluation, based on
  Berkeley’s Roofline model demonstrates that our implementation is near op
 timally tuned on the Xilinx Alveo U280.\n\nTag: Accelerator-based Architec
 tures, Applications, Architectures, Emerging Technologies, Heterogeneous S
 ystems, Memory Systems, Networks\n\nRegistration Category: Workshop Reg Pa
 ss
END:VEVENT
END:VCALENDAR
