BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055401Z
LOCATION:Second Floor Atrium
DTSTART;TZID=America/Chicago:20211118T083000
DTEND;TZID=America/Chicago:20211118T170000
UID:submissions.supercomputing.org_SC21_sess280_rpost144@linklings.com
SUMMARY:Padding to Extend the Bruck Algorithm for Non-Uniform All-to-All C
 ommunication
DESCRIPTION:Posters, Research Posters\n\nPadding to Extend the Bruck Algor
 ithm for Non-Uniform All-to-All Communication\n\nFan, Gilray, Kumar\n\nThe
  latency of the standard MPI_Alltoallv implementations is linear in the nu
 mber of processes. Such linear complexity performs poorly when application
 s are deployed on millions of cores for short messages, which is dominated
  by latency. Bruck's algorithm is a classic logarithm algorithm for unifor
 m all-to-all communication. It fails, however, to support messages of vary
 ing sizes. In this paper, we present the Padded Bruck algorithm, a natural
  extension strategy for applying Bruck’s algorithm to non-uniform all-to-a
 ll communication by transforming it into uniform all-to-all communication.
  We also analyze several variants of Bruck’s algorithm and investigate the
  underlying causes of their behavior, with the ultimate goal of gaining in
 sights that can be applied to our Padded Bruck algorithm. When compared to
  Cray’s MPI_Alltoallv, our evaluation shows that our algorithm outperforms
  in most cases if the message size is less than 1024 bytes and the process
  count is smaller than 8192.\n\nRegistration Category: Tech Program Reg Pa
 ss, Exhibit Hall Only
END:VEVENT
END:VCALENDAR
