BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055339Z
LOCATION:Online
DTSTART;TZID=America/Chicago:20211114T160500
DTEND;TZID=America/Chicago:20211114T163000
UID:submissions.supercomputing.org_SC21_sess428_ws_ftxs108@linklings.com
SUMMARY:Relaxed Replication for Energy Efficient and Resilient GPU Computi
 ng
DESCRIPTION:Workshop\n\nRelaxed Replication for Energy Efficient and Resil
 ient GPU Computing\n\nMiao, Calhoun, Ge\n\nPower and reliability are two  
 intertwined challenges in GPU-accelerated large-scale computing. Aggressiv
 e power reduction pushes hardware to its operating limit and increases the
  failure rate. Resilience allows programs to progress when subjected to fa
 ults and is an integral component of large-scale systems, but incurs signi
 ficant time and energy overhead.  Managing power and resilience is challen
 ging, due to the heterogeneous compute capability, power consumption, and 
 varying failure rates between CPUs and GPUs. Previous works have shown tha
 t redundancy-based approaches are more energy efficient than checkpointing
 /restart at extreme-scales, but current solutions only support parallel pr
 ograms running on CPU-based homogeneous systems. Simply extending redundan
 cy approaches from CPU-based systems results in sub-optimal performance an
 d/or energy efficiency because existing redundancy solutions typically rel
 y on identical replicas with expensive synchronization. In this work, we e
 xplore redundancy techniques and energy efficient techniques for GPU-accel
 erated systems running MPI parallel workloads. Specifically, we design a n
 ovel redundancy technique that relaxes the requirement of synchronization 
 and identicalness for replica processes and allows them to run in lower-pr
 ecision  and at lower power/performance states with periodical rejuvenatio
 n or asynchronization, enabling resources and power reduction. \n\nThis re
 laxed replication mechanism complicates fault detection and recovery over 
 the homogeneous exact replication. We discuss techniques to handle and mit
 igate these complexities for both process/node failures and silent data co
 rruption. Evaluation results on  a 16-GPU cluster show our techniques redu
 ce energy by up to 15% for unmodified programs and 32% for programs that a
 re able to adapt the precision of the replicas.\n\nTag: Online Only, Extre
 me Scale Computing, Reliability and Resiliency\n\nRegistration Category: W
 orkshop Reg Pass
END:VEVENT
END:VCALENDAR
