BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055339Z
LOCATION:226
DTSTART;TZID=America/Chicago:20211114T145300
DTEND;TZID=America/Chicago:20211114T150000
UID:submissions.supercomputing.org_SC21_sess431_ws_llvmlt108@linklings.com
SUMMARY:Remote OpenMP Offloading
DESCRIPTION:Workshop\n\nRemote OpenMP Offloading\n\nPatel, Doerfert\n\nIn 
 this work we show that the OpenMP accelerator offloading model is sufficie
 nt to seamlessly and efficiently utilize more than a single compute node, 
 and its connected accelerators.  \n\nWithout source code or compiler modif
 ications we run an OpenMP offload capable program on a remote CPU, or remo
 te accelerator (e.g., GPU), as if it was a local one.  For applications th
 at support multi-device offloading, any combination of local and remote CP
 Us and accelerators can be utilized simultaneously, fully transparent to t
 he user.  Our low-overhead implementation is integrated into the LLVM/Open
 MP compiler infrastructure as a plugin and is publicly available (in parts
 ) with LLVM 12 and later.\n\nTo evaluate our work we provide detailed stud
 ies on scaling results for two HPC proxy applications. We show perfect sca
 ling across dozens of GPUs in multiple hosts with effectiveness proportion
 al to the ratio of computation versus memory transfer time.\n\nTag: Parall
 el Programming Systems\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
