BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055405Z
LOCATION:Online
DTSTART;TZID=America/Chicago:20211119T103000
DTEND;TZID=America/Chicago:20211119T105000
UID:submissions.supercomputing.org_SC21_sess334_ws_lasalss108@linklings.co
 m
SUMMARY:Unleashing the Performance of bmSparse for the Sparse Matrix Multi
 plication in GPUs
DESCRIPTION:Workshop\n\nUnleashing the Performance of bmSparse for the Spa
 rse Matrix Multiplication in GPUs\n\nBerger, Freire, Marini, Dufrechou, Ez
 zatti\n\nThe evolution of data science and machine learning has increased 
 the applicability of the sparse matrix multiplication (SPGEMM) kernel. Unl
 ike more well-known operations such as the SPMV, in the SPGEMM the nonzero
  pattern of the result is determined by the interaction between the nonzer
 o patterns of the inputs, which imposes serious challenges to the developm
 ent of high-performance implementations for accelerators. Recent efforts i
 n this subject aim to mitigate this irregularity through the use of block-
 based sparse storage formats, obtaining promising results on accelerators 
 such as GPUs. In this work, we study the format bmSparse [1] and propose o
 ptimizations to attack the principal bottlenecks of the original SPGEMM im
 plementation for Nvidia GPUs. We evaluate the proposal using nine sparse m
 atrices of different sizes, showing remarkable speedups with respect to CU
 SPARSE’s CSR variant.\n\nTag: Online Only, Algorithms, Extreme Scale Compu
 ting\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
