BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055339Z
LOCATION:Online
DTSTART;TZID=America/Chicago:20211114T113500
DTEND;TZID=America/Chicago:20211114T120000
UID:submissions.supercomputing.org_SC21_sess428_ws_ftxs101@linklings.com
SUMMARY:Statistical Framework for Two-Party Acceptance Testing of HPC Syst
 ems for Reliability
DESCRIPTION:Workshop\n\nStatistical Framework for Two-Party Acceptance Tes
 ting of HPC Systems for Reliability\n\nDeBardeleben, Burr, Penton, Walker,
  Loncaric...\n\nHPC clusters and supercomputers are capital investments an
 d undergo great scrutiny to be sure that the system meets agreed upon metr
 ics of performance, reliability, and usability.  As such, careful evaluati
 on of a system occurs once delivered to evaluate the agreed upon specifica
 tions have been met.  This evaluation is referred to as an "acceptance tes
 t." Both the HPC vendor and the data center buying the system have a veste
 d interest in passing this test, though their goals sometimes are at odds.
   While the buyer wants the system agreed upon, the vendor wants the test 
 to pass in a timely manner so that they can be paid.  This creates a delic
 ate balance where both parties agree upon testing parameters to satisfy th
 eir goals and optimize their objectives.\n\nThis paper focuses on the reli
 ability testing aspect of an acceptance test and outlines how test paramet
 ers are set up for length of test and number of acceptable failures.  Seve
 ral statistical approaches are presented that illuminate the relationships
  among salient model parameters and both vendor and buyer constraints. Add
 itionally, techniques for accepting a less reliable machine and those rami
 fications are presented as well. Finally, simulations are performed that a
 nalyze six different HPC workloads and how impactful accepting a less reli
 able machine will be on the buyer.  The techniques presented in this paper
  can be used by data center operators and procurement teams to evaluate sy
 stems during acceptance testing as well as HPC vendors to minimize their r
 isk of failure.\n\nTag: Online Only, Extreme Scale Computing, Reliability 
 and Resiliency\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
