BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055339Z
LOCATION:226
DTSTART;TZID=America/Chicago:20211114T115000
DTEND;TZID=America/Chicago:20211114T123000
UID:submissions.supercomputing.org_SC21_sess431_ws_llvmf102@linklings.com
SUMMARY:Extending LLVM IR for DPC++ Matrix Support: A Case Study with Inte
 lŽ Advanced Matrix Extensions (IntelŽ AMX)
DESCRIPTION:Workshop\n\nExtending LLVM IR for DPC++ Matrix Support: A Case
  Study with IntelŽ Advanced Matrix Extensions (IntelŽ AMX)\n\nKhaldi, Luo,
  Yu, Sotkin, Morais...\n\nIn this paper, we introduce a DPC++ matrix ex-te
 nsion to unify different tensor hardware: IntelŽ Advanced Matrix Extension
 s (IntelŽ AMX) to CPUs, NVIDIAŽ TPUs, IBMŽ POWERŽ MMA, etc.  These tensor 
 hardware units are usually accessed by low-level intrinsics or assembly to
  perform matrix operations.  It is hard for scientists to program these do
 main-specific devices without the kind of high-level abstractions and effi
 cient implementations we introduce here.  We also extend the existing LLVM
  matrix intrinsics to represent this DPC++ extension and yield efficient I
 ntel AMX code generation.  Based on our case study of implementing this in
 terface  on  Intel AMX  hardware, we discuss some of the limitations of ex
 isting LLVM Intermediate Representation (IR) and how they can be overcome 
 to exploit tensor hardware.\n\nTag: Parallel Programming Systems\n\nRegist
 ration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
