BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055411Z
LOCATION:Online
DTSTART;TZID=America/Chicago:20211115T080000
DTEND;TZID=America/Chicago:20211115T170000
UID:submissions.supercomputing.org_SC21_sess208_tut131@linklings.com
SUMMARY:Node-Level Performance Engineering
DESCRIPTION:Tutorial\n\nNode-Level Performance Engineering\n\nHager, Eitzi
 nger, Wellein\n\nAs we move towards exascale, the gap between peak and app
 lication performance continues to open. Paradoxically, slow code tends to 
 be highly scalable. Consequently, valuable resources are wasted, often on 
 a massive scale. If the user values resource efficiency on any scale, opti
 mal performance on the node level is paramount. We convey the architectura
 l features of current processor chips, multiprocessor nodes, and accelerat
 ors, as far as they are relevant for the practitioner. Peculiarities like 
 SIMD, cache topology, bandwidth bottlenecks, and ccNUMA characteristics ar
 e introduced, and the influence of system topology and affinity on the per
 formance of parallel code is demonstrated. Performance engineering is intr
 oduced as a powerful tool that helps the user understand the bottlenecks a
 t hand and to assess the impact of optimizations. A cornerstone of these c
 oncepts is the roofline model, which is described in detail with useful ca
 se studies and limits of its applicability.\n\nTag: Online Only, Performan
 ce\n\nRegistration Category: Tutorial Reg Pass
END:VEVENT
END:VCALENDAR
