BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055400Z
LOCATION:220-221
DTSTART;TZID=America/Chicago:20211119T105500
DTEND;TZID=America/Chicago:20211119T112000
UID:submissions.supercomputing.org_SC21_sess339_misc391@linklings.com
SUMMARY:Survival of the Fittest Amidst the Cambrian Explosion of Processor
  Architectures for Artificial Intelligence
DESCRIPTION:Workshop\n\nSurvival of the Fittest Amidst the Cambrian Explos
 ion of Processor Architectures for Artificial Intelligence\n\nSukumar\n\nT
 he need for high performance computing in data-driven artificial intellige
 nce (AI) workloads has led to the Cambrian explosion of processor architec
 tures. As these novel processor architectures aim to evolve and thrive ins
 ide datacenters and cloud-services, we need to understand different figure
 s-of-merit for device-, server- and rack- scale systems. Towards that goal
 , we share early-access hands-on experience with these processor/accelerat
 or architectures. We describe an evaluation plan that includes carefully c
 hosen neural network models to gauge the maturity of the hardware and soft
 ware ecosystem. Our hands-on evaluation using benchmarks reveals significa
 nt benefits of hardware acceleration while exposing several blind spots in
  the software ecosystem. Ranking the benefits based on different figures o
 f merit such as cost, energy, and adoption efficiency reveals a “heterogen
 ous” future for production systems with multiple processor architectures i
 n the edge-to-datacenter AI workflow. \n\nPreparing to survive in this he
 terogenous future, we describe a method to profile and predict the perform
 ance benefits of a deep learning training workload on novel architectures.
  Our approach profiles the neural network model for memory, bandwidth and 
 compute requirements by analyzing the model definition. Then, using profil
 ing tools, we estimate the I/O and arithmetic intensity requirements at di
 fferent batch sizes. By overlaying profiler results onto analytic roofline
  models of the emerging processor architectures, we identify opportunities
  for potential acceleration. We discuss how the interpretation of the roof
 line analysis can guide system architecture to deliver productive performa
 nce and conclude with recommendations to survive the Cambrian explosion.\n
 \nTag: Heterogeneous Systems, Parallel Programming Languages and Models, P
 arallel Programming Systems, Productivity Tools, System Software and Runti
 me Systems\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
