BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055402Z
LOCATION:231-232
DTSTART;TZID=America/Chicago:20211115T170000
DTEND;TZID=America/Chicago:20211115T173000
UID:submissions.supercomputing.org_SC21_sess343_ws_h2rc107@linklings.com
SUMMARY:Porting In-Compressible Flow Matrix Assembly to FPGAs for Accelera
 ting HPC Engineering Simulations
DESCRIPTION:Workshop\n\nPorting In-Compressible Flow Matrix Assembly to FP
 GAs for Accelerating HPC Engineering Simulations\n\nBrown\n\nEngineering i
 s an important domain for supercomputing, with the Alya model being a popu
 lar code for undertaking such simulations. With ever increasing demand fro
 m users to model larger, more complex systems at reduced time to solution 
 it is important to explore the role that novel hardware technologies, such
  as FPGAs, can play in accelerating these workloads on future exascale sys
 tems. \n\nIn this paper, we explore the porting of Alya's in-compressible 
 flow matrix assembly kernel, which accounts for a large proportion of the 
 model runtime, onto FPGAs. After describing in detail successful strategie
 s for optimisation at the kernel level, we then explore sharing the worklo
 ad between the FPGA and host CPU, mapping most appropriate parts of the ke
 rnel between these technologies, enabling us to more effectively exploit t
 he FPGA. We then compare the performance of our approach on a Xilinx Alveo
  U280 against a 24-core Xeon Platinum CPU and Nvidia V100 GPU, with the FP
 GA significantly out-performing the CPU and performing comparably against 
 the GPU, whilst drawing substantially less power. The result of this work 
 is both an experience report describing appropriate dataflow optimisations
  which we believe can be applied more widely across HPC codes, and a perfo
 rmance comparison for this specific workload that demonstrates the potenti
 al for FPGAs in accelerating HPC engineering simulations.\n\nTag: Accelera
 tor-based Architectures, Applications, Architectures, Emerging Technologie
 s, Heterogeneous Systems, Memory Systems, Networks\n\nRegistration Categor
 y: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
