BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055339Z
LOCATION:Online
DTSTART;TZID=America/Chicago:20211114T090000
DTEND;TZID=America/Chicago:20211114T173000
UID:submissions.supercomputing.org_SC21_sess428@linklings.com
SUMMARY:FTXS: Workshop on Fault-Tolerance for HPC at Extreme Scale
DESCRIPTION:Workshop\n\nCharacterizing Per-Node Memory Failures Using Benf
 ord’s Law\n\nFerreira, Levy\n\nFault tolerance is a key challenge as high 
 performance computing systems continue to increase component counts, indiv
 idual component reliability decreases, and hardware and software complexit
 y increases. To better understand the potential impacts of failures on nex
 t-generation systems, significant e...\n\n---------------------\nAccelerat
 ing Checkpoint/Restart with Lossy Methods\n\nIldes, Kastoras, Keller, Baut
 ista Gomez\n\nApproximate computing targets applications with the ability 
 to tolerate losses of accuracy in the computational results. The essence o
 f approximate computing is to use a data representation, that allows to re
 duce the data size or speed up computations at the cost of data accuracy. 
 \n\nData reduction i...\n\n---------------------\nFTXS 2021 : Oren et al. 
 Q&A\n\nFridman\n\nQ&A following " Assessing the Use Cases of Persistent Me
 mory in High-Performance Scientific Computing" Oren, Fridman.\n\n---------
 ------------\nDoubt and Redundancy Kill Soft Errors—Toward Detection and C
 orrection of Silent Data Corruption in Task-Based Numerical Software\n\nSa
 mfass, Weinzierl, Reinarz, Bader\n\nResilient algorithms in high-performan
 ce computing are subject to rigorous non-functional constraints. Resilienc
 y must not increase the runtime, memory footprint or I/O demands too signi
 ficantly. We propose a task-based soft error detection scheme that relies 
 on error criteria per task outcome. They...\n\n---------------------\nInco
 rporating Fault-Tolerance Awareness into System-Level Modeling and Simulat
 ion\n\nJohnson, Lam\n\nAs the design space for supercomputers grows, model
 ing and simulation (MODSIM) becomes more important to facilitate system-le
 vel design space exploration (DSE). Furthermore, extreme-scale systems and
  newer technologies can lead to higher fault rates, which negatively affec
 ts system performance. Ther...\n\n---------------------\nFTXS Lunch Break 
 (12:30-2:00)\n\n\n\n---------------------\nFTXS 2021 : Ildes et al. Q&A\n\
 nIldes, Kastoras\n\nQ&A following "Accelerating checkpoint/restart with lo
 ssy methods", Ildes, Kastoras, Keller, Bautista Gomez.\n\n----------------
 -----\nFTXS Afternoon Break (3:00-3:30pm)\n\n\n\n---------------------\nFT
 XS Morning Break (10-10:30)\n\n\n\n---------------------\nFTXS 2021 : Miao
  Q&A\n\nMiao\n\nQ&A following "Relaxed Replication for Energy Efficient an
 d Resilient GPU Computing", Miao.\n\n---------------------\nAssessing the 
 Use Cases of Persistent Memory in High-Performance Scientific Computing\n\
 nOren, Fridman\n\nAs the High Performance Computing world moves towards th
 e Exa-Scale era, huge amounts of data should be analyzed, manipulated and 
 stored. In the traditional storage/memory hierarchy, whenever the DRAM's c
 apacity becomes insufficient for storing data in a node, the computation s
 hould either be distri...\n\n---------------------\nFTXS 2021 : Johnson et
  al. Q&A\n\nJohnson\n\nQ&A following "Incorporating Fault-Tolerance Awaren
 ess into System-Level Modeling and Simulation", Johnson, Lam.\n\n---------
 ------------\nRelaxed Replication for Energy Efficient and Resilient GPU C
 omputing\n\nMiao, Calhoun, Ge\n\nPower and reliability are two  intertwine
 d challenges in GPU-accelerated large-scale computing. Aggressive power re
 duction pushes hardware to its operating limit and increases the failure r
 ate. Resilience allows programs to progress when subjected to faults and i
 s an integral component of large-scal...\n\n---------------------\nFTXS: W
 orkshop on Fault-Tolerance for HPC at Extreme Scale\n\nLevy, Teranishi, Da
 ly\n\nIncreases in the number, variety and complexity of components requir
 ed to compose next-generation extreme-scale systems mean that systems will
  experience significant increases in aggregate fault rates, fault diversit
 y and fault complexity. Additionally, the widespread availability of new s
 torage dev...\n\n---------------------\nFTXS 2021 : Featured Speaker - Dr.
  Catherine Schuman (Fault Tolerance and Resilience in Neuromorphic Systems
 )\n\nSchuman\n\nBrain-inspired neuromorphic computing systems are promisin
 g for the future of computing beyond the looming end of Moore's law.  Thou
 gh robustness and resilience are commonly quoted as features of neuromorph
 ic computing systems, the expected performance of neuromorphic systems in 
 the face of hardware...\n\n---------------------\nFTXS 2021 : Featured Spe
 aker Q&A\n\nSchuman\n\nQ&A following the featured speaker.\n\n------------
 ---------\nFTXS 2021 : Opening Remarks\n\nLevy\n\nWelcome to FTXS 2021!\n\
 n---------------------\nFTXS 2021 : Ferreira et al. Q&A\n\nFerreira\n\nQ&A
  following "Characterizing Per-node Memory Failures Using Benford’s Law", 
 Ferreira, Levy.\n\n---------------------\nStatistical Framework for Two-Pa
 rty Acceptance Testing of HPC Systems for Reliability\n\nDeBardeleben, Bur
 r, Penton, Walker, Loncaric...\n\nHPC clusters and supercomputers are capi
 tal investments and undergo great scrutiny to be sure that the system meet
 s agreed upon metrics of performance, reliability, and usability.  As such
 , careful evaluation of a system occurs once delivered to evaluate the agr
 eed upon specifications have been met....\n\n---------------------\nFTXS 2
 021 : Closing Remarks\n\nLevy\n\n---------------------\nFTXS 2021 : Samfas
 s et al. Q&A\n\nSamfass\n\nQ&A following "Doubt and Redundancy Kill Soft E
 rrors—Towards Detection and Correction of Silent Data Corruption in Task-b
 ased Numerical Software", Samfass, Weinzierl, Reinarz, Bader.\n\n---------
 ------------\nFTXS 2021 : Debardeleben et al. Q&A\n\nDebardeleben\n\nQ&A f
 ollowing "Statistical Framework for Two-Party Acceptance Testing of HPC Sy
 stems for Reliability", DeBardeleben, Burr, Penton, Walker, Loncaric, Jone
 s.\n\n---------------------\nFTXS 2021: Invited Speaker - Devesh Tiwari (M
 aking Erroneous Executions on Quantum Computers Meaningful)\n\nTiwari\n\nQ
 uantum computing has moved from theoretical promise to practical realizati
 on -- at a very rapid pace in the last decade. But, prohibitively high err
 or rates on existing Near-term Intermediate-Scale Quantum (NISQ) computers
  limit their usability even for quantum-advantage-proven algorithms (that 
 is,...\n\n\nTag: Online Only, Extreme Scale Computing, Reliability and Res
 iliency\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
