BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T054757Z
LOCATION:230-231-232
DTSTART;TZID=America/Chicago:20211116T110000
DTEND;TZID=America/Chicago:20211116T113000
UID:submissions.supercomputing.org_SC21_sess177_pap467@linklings.com
SUMMARY:Tensor Processing Primitives: A Programming Abstraction for Effici
 ency and Portability in Deep Learning Workloads
DESCRIPTION:Paper\n\nTensor Processing Primitives: A Programming Abstracti
 on for Efficiency and Portability in Deep Learning Workloads\n\nGeorganas,
  Kalamkar, Avancha, Adelman, Anderson...\n\nDuring the past decade, novel 
 deep learning (DL) algorithms/workloads and hardware have been developed t
 o tackle a wide range of problems. Despite the advances in workload/hardwa
 re ecosystems, the programming methodology of DL-systems is stagnant. DL-w
 orkloads leverage either highly-optimized, yet platform-specific and infle
 xible, kernels from DL-libraries, or as for novel operators, reference imp
 lementations are built via DL-framework primitives with underwhelming perf
 ormance. This work introduces the Tensor Processing Primitives (TPP), a pr
 ogramming abstraction striving for efficient, portable implementation of D
 L-workloads with high productivity. TPPs define a compact, yet versatile s
 et of 2D-tensor operators, which subsequently can be utilized as building-
 blocks to construct complex operators on high-dimensional tensors. The TPP
  specification is platform-agnostic, thus code expressed via TPPs is porta
 ble, whereas the TPP implementation is highly-optimized and platform-speci
 fic. We demonstrate the efficacy of our approach using standalone kernels 
 and end-to-end DL-workloads expressed entirely via TPPs that outperform st
 ate-of-the-art implementations on multiple platforms.\n\nTag: Reproducibil
 ity Badge, Machine Learning and Artificial Intelligence\n\nRegistration Ca
 tegory: Tech Program Reg Pass\n\nReproducibility Badges: Artifact Availabl
 e, Artifact Functional, Results Reproduced
END:VEVENT
END:VCALENDAR
