BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055342Z
LOCATION:227
DTSTART;TZID=America/Chicago:20211115T103000
DTEND;TZID=America/Chicago:20211115T110000
UID:submissions.supercomputing.org_SC21_sess347_ws_scsc107@linklings.com
SUMMARY:Toward Aggregated Asynchronous Checkpointing
DESCRIPTION:Workshop\n\nToward Aggregated Asynchronous Checkpointing\n\nGo
 ssman\n\nHigh-Performance Computing (HPC) applications need to check-point
  massive data sizes at scale with increasing frequency. Multi-level asynch
 ronous checkpoint runtimes like VELOC (Very Low Overhead Checkpoint Strate
 gy) are gaining popularity among application scientists for their ability 
 to leverage fast node-local storage and flush independently to stable, ext
 ernal storage (e.g., parallel file systems) in the background. Currently, 
 VELOC adopts a one-file-per-process flush strategy, which results in a lar
 ge number of files being written to external storage, thereby overwhelming
  metadata servers and making it difficult to transfer and access checkpoin
 ts as a whole. This paper discusses the challenges and opportunities of de
 signing aggregation techniques for asynchronous multi-level checkpointing.
  To this end, we implement and studied two aggregation strategy, study the
 ir limitations and propose a new aggregation strategy specifically for asy
 nchronous multi-level checkpointing.\n\nTag: Parallel Programming Language
 s and Models, Reliability and Resiliency\n\nRegistration Category: Worksho
 p Reg Pass
END:VEVENT
END:VCALENDAR
