BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055412Z
LOCATION:Online
DTSTART;TZID=America/Chicago:20211114T160500
DTEND;TZID=America/Chicago:20211114T161000
UID:submissions.supercomputing.org_SC21_sess328_ws_whpc108@linklings.com
SUMMARY:Early Career Lighting Talks – Atos: A Task-Parallel GPU Dynamic Sc
 heduling Framework for Dynamic Irregular Computations
DESCRIPTION:Workshop\n\nEarly Career Lighting Talks – Atos: A Task-Paralle
 l GPU Dynamic Scheduling Framework for Dynamic Irregular Computations\n\nC
 hen\n\nWe present Atos, a task-parallel GPU dynamic scheduling framework t
 hat is especially suited to dynamic irregular applications. Compared to th
 e dominant Bulk Synchronous Parallel (BSP) frameworks, Atos removes the gl
 obal barriers and exposes additional concurrency by supporting task-parall
 el formulations of applications with relaxed dependencies, achieving highe
 r GPU utilization, which is particularly significant for problems with con
 currency bottlenecks. Atos also offers implicit task-parallel load balanci
 ng in addition to data-parallel load balancing, providing users the flexib
 ility to balance between them to achieve optimal performance. Finally, Ato
 s allows users to adapt to different use cases by controlling the kernel s
 trategy and task-parallel granularity. We demonstrate that each of these c
 ontrols is important in practice.\n\nWe evaluate and analyze the performan
 ce of Atos vs. BSP on two applications: breadth-first search and PageRank.
  Atos implementations achieve geomean speedups of {3.44x, 2.1x} and peak s
 peedups of {12.8x, 3.2x} on two case studies respectively, compared to a s
 tate-of-the-art BSP GPU implementation.\n\nIn the future, we plan to exten
 d this framework to multi-GPU and multi-node systems and expect Atos’s tas
 k-based, global-synchronization-free programming model is likely to be mor
 e amenable for use in a distributed environment. As the number of GPUs in 
 modern computer systems increases, the fraction of total runtime spent on 
 communication and synchronization also increases. Under distributed contex
 t, Atos 1) removes the global barriers thus largely reducing the synchroni
 zation cost; 2) is able to generate fine-grained messages and send them im
 mediately without synchronization, leading to better communication-computa
 tion overlap.\n\nTag: Online Only, Career Development, Diversity Equity In
 clusion (DEI), Education and Training and Outreach, HPC Community Collabor
 ation\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
