BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055339Z
LOCATION:Online
DTSTART;TZID=America/Chicago:20211114T122000
DTEND;TZID=America/Chicago:20211114T123000
UID:submissions.supercomputing.org_SC21_sess428_ws_ftxs107@linklings.com
SUMMARY:Characterizing Per-Node Memory Failures Using Benford’s Law
DESCRIPTION:Workshop\n\nCharacterizing Per-Node Memory Failures Using Benf
 ord’s Law\n\nFerreira, Levy\n\nFault tolerance is a key challenge as high 
 performance computing systems continue to increase component counts, indiv
 idual component reliability decreases, and hardware and software complexit
 y increases. To better understand the potential impacts of failures on nex
 t-generation systems, significant effort has been devoted to collecting, c
 haracterizing and analyzing failures on current systems. These studies req
 uire large volumes of data and complex analysis in an attempt to identify 
 statistical properties of the failure data. In this paper, we examine the 
 lifetime of failures on the Cielo supercomputer that was located at Los Al
 amos National Laboratory, looking specifically at the per-node time betwee
 n faults. Through this analysis, we show that the time between correctable
  faults on nodes obeys Benford’s law, This law applies to a number of natu
 rally occurring collections of numbers and states that the leading digit i
 s more likely to be small, for example a leading digit of 1 is more likely
  than 9. This is in contrast to previous work that examined the interarriv
 al time for correctable faults on the entire machine, which do not obey Be
 nford’s law. This initial work provides critical analysis on the distribut
 ion of times between failures for extreme-scale systems. More specifically
 , the distributed analysis technique outlined in this work has the potenti
 al for significantly lower overheads than the centralized approach describ
 ed in previous work. Also, this work enables a simple form of distributed 
 failure prediction that can be utilized to lower failure mitigation overhe
 ads.\n\nTag: Online Only, Extreme Scale Computing, Reliability and Resilie
 ncy\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
