BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055412Z
LOCATION:Online
DTSTART;TZID=America/Chicago:20211114T143000
DTEND;TZID=America/Chicago:20211114T150000
UID:submissions.supercomputing.org_SC21_sess429_ws_p3hpc110@linklings.com
SUMMARY:Case Study of Using Kokkos and SYCL as Performance-Portable Framew
 orks for MILC-DSLASH Benchmark on NVIDIA, AMD, and Intel GPUs
DESCRIPTION:Workshop\n\nCase Study of Using Kokkos and SYCL as Performance
 -Portable Frameworks for MILC-DSLASH Benchmark on NVIDIA, AMD, and Intel G
 PUs\n\nDufek, Gayatri, Mehta, Doerfler, Cook...\n\nIn this paper, we intro
 duce a GPU-friendly parallel implementation of Milc-Dslash that exposes mu
 ltiple hierarchies of parallelism in the algorithm. Milc-Dslash was design
 ed to serve as a benchmark with highly optimized matrix-vector multiplicat
 ions to measure the resource utilization on the GPU systems. The parallel 
 hierarchies in the Milc-Dslash algorithm are mapped onto a target hardware
  using Kokkos and SYCL programming models. We present the performance achi
 eved by Kokkos and SYCL implementations of Milc-Dslash on NVIDIA A100 GPU,
  AMD MI100 GPU, and Intel Gen9 GPU. Additionally, we compare the Kokkos an
 d SYCL performances with those obtained from the versions written in CUDA 
 and HIP programming models on NVIDIA A100 GPU and AMD MI100 GPU, respectiv
 ely.\n\nTag: Online Only, Heterogeneous Systems, Parallel Programming Lang
 uages and Models, Performance, Productivity Tools, Software Engineering\n\
 nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
