BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055346Z
LOCATION:Online
DTSTART;TZID=America/Chicago:20211115T090000
DTEND;TZID=America/Chicago:20211115T173000
UID:submissions.supercomputing.org_SC21_sess422@linklings.com
SUMMARY:PMBS21: The 12th International Workshop on Performance Modeling, B
 enchmarking and Simulation of High-Performance Computer Systems
DESCRIPTION:Workshop\n\nUsing the Semi-Stencil Algorithm to Accelerate Hig
 h-Order Stencils on GPUs\n\nSai, Mellor-Crummey, Meng, Araya-Polo, Meng\n\
 nUnderstanding how to develop efficient high-order stencils for graphics p
 rocessing units (GPUs) is a topic of great interest for many application d
 omains. High-performance stencils on GPUs must be tailored for data parall
 el computation and to use the memory hierarchy efficiently. For data-inten
 sive ...\n\n---------------------\nPMBS21:  Introduction and Welcome\n\n\n
 \n---------------------\nBayesian Optimization for Auto-Tuning GPU kernels
 \n\nWillemsen, van Nieuwpoort, van Werkhoven\n\nFinding optimal parameter 
 configurations for tunable GPU kernels is a non-trivial exercise for large
  search spaces, even when automated. This poses an optimization task on a 
 non-convex search space, using an expensive-to-evaluate function with unkn
 own derivative.  These characteristics make a good c...\n\n---------------
 ------\nMemory Demands in Disaggregated HPC: How Accurate Do We Need to Be
 ?\n\nVieira Zacarias, Petrucci, Carpenter\n\nJobs running on HPC systems c
 an vary dramatically due to the intrinsic differences in application resou
 rce requirements (e.g. memory or cores). Since HPC applications run on a n
 umber of self-contained servers whose capacities are fixed at design time,
  there is often a mismatch between the resource p...\n\n------------------
 ---\nCustomized Monte Carlo Tree Search for LLVM/Polly's Composable Loop O
 ptimization Transformations\n\nKoo, Balaprakash, Kruse, Wu, Hovland...\n\n
 Polly is the LLVM project's polyhedral loop optimizer. Recent user-directe
 d loop transformation pragmas were proposed based on LLVM/Clang and Polly.
  The search space exposed by the transformation pragmas is a tree, wherein
  each node represents a specific combination of loop transformations. To f
 ind ...\n\n---------------------\nArchitectural Requirements for Deep Lear
 ning Workloads in HPC Environments\n\nIbrahim, Nguyen, Nam, Bhimji, Farrel
 l...\n\nScientific machine learning (SciML) promises to have a transformat
 ional impact on scientific exploration, by combining state-of-the-art AI m
 ethods with the latest generation of supercomputers. To efficiently levera
 ge ML techniques on high-performance computing (HPC) systems, however, it 
 is critical ...\n\n---------------------\nPMBS21:  Afternoon Break (3-3:30
 )\n\n\n\n---------------------\nExploration of Congestion Control Techniqu
 es on Dragonfly-Class HPC Networks Through Simulation\n\nMcGlohon, Hemmert
 , Brown, Levenhagen, Chunduri...\n\nEnsuring optimal communication latency
  in high-performance computing (HPC) networks is of critical importance to
  the efficient operation of facilitated applications. Different applicatio
 n operations and types of tasks, such as I/O operations, can create a vari
 ety of traffic patterns across the syste...\n\n---------------------\nMult
 ilevel Simulation-Based Co-Design of Next Generation HPC Microprocessors\n
 \nZaourar, Benazouz, Mouhagir, Jebali, Sassolas...\n\nThis paper demonstra
 tes the combined use of three simulation tools in support of a full co-des
 ign methodology for an HPC-focused SoC. The simulation tools make differen
 t trade-offs among simulation speed, accuracy and model abstraction level,
  and are shown to be complementary to one another. We appl...\n\n---------
 ------------\nComparing Julia to Performance Portable Parallel Programming
  Models for HPC\n\nLin, McIntosh-Smith\n\nJulia is a general-purpose, mana
 ged, strongly and dynamically-typed programming language with emphasis on 
 high-performance scientific computing. Traditionally, HPC software develop
 ment uses languages such as C, C++ and Fortran, which compile to unmanaged
  code. This offers the programmer near bare-me...\n\n---------------------
 \nPMBS21:  Lunch Break (12:30-2)\n\n\n\n---------------------\nMicroBench 
 Maker: Reproduce, Reuse, Improve\n\nHunold, Ajanohoun, Carpen-Amarie\n\nBe
 nchmarking is one of the fundamental methods for analyzing the performance
  of computational processes or threads. In the domain of high-performance 
 computing, benchmarks are essential to assess computer systems; e.g., the 
 TOP500 or the Green500 benchmarks are used to define the performance of ma
 ch...\n\n---------------------\nPMBS21: The 12th International Workshop on
  Performance Modeling, Benchmarking and Simulation of High-Performance Com
 puter Systems\n\nWright, Jarvis, Hammond\n\nThe PMBS21 workshop is concern
 ed with the comparison of high-performance computing systems through perfo
 rmance modeling, benchmarking or through the use of tools such as simulato
 rs. We are particularly interested in research which reports the ability t
 o measure and make tradeoffs in software/hardwar...\n\n-------------------
 --\nNarrowing the Search Space of Applications Mapping on Hierarchical Top
 ologies\n\nDenoyelle, Jeannot, Perarnau, Videau, Beckman\n\nProcessor arch
 itectures at exascale and beyond are expected to continue to suffer from i
 ssues of nonuniform access to in-die and node-wide shared resources. Mappi
 ng applications onto these resource hierarchies is an on-going performance
  concern, requiring specific care for increasing locality and re...\n\n---
 ------------------\nUnderstanding Power Variation and its Implications on 
 Performance Optimization on the Cori Supercomputer\n\nBhalachandra, Austin
 , Wright\n\nPower is increasingly becoming a limiting factor in supercompu
 ting.  The performance and scale of future high-performance computing syst
 ems will be determined by how efficiently they manage their power budgets.
  Therefore, any amount of unused power is forsaken performance. Regardless
  of the processo...\n\n---------------------\nPMBS21:  Morning Break (10-1
 0:30)\n\n\n\n---------------------\nAn Extended Roofline Performance Model
  with PCI-E and Network Ceilings\n\nDufek, Deslippe, Lin, Yang, Cook...\n\
 nIn this work, we evaluate the utility of adding two new diagonal ceilings
  to the roofline model related to PCI-E and effective network bandwidths t
 o provide insights into how communication impacts the performance of large
 -scale parallel applications. The roofline performance analysis is based o
 n two...\n\n---------------------\nEnabling Cache-Aware Roofline Analysis 
 with Portable Hardware Counter Metrics\n\nGravelle, Nystrom, Yokelson, Nor
 ris\n\nIn this paper,  we seek to provide guidance on how to empirically c
 ollect the information required to plot application points on Intel’s Casc
 ade Lake Xeon processor and Fujitsu’s  A64FX  ARM-based  processor.  Under
 standing how to process this information is vital to use the Roofline for 
 empirical p...\n\n\nTag: Online Only, Accelerator-based Architectures, App
 lications, Computational Science, Emerging Technologies, Extreme Scale Com
 puting, File Systems and I/O, Heterogeneous Systems, Parallel Programming 
 Languages and Models, Performance, Scientific Computing, Software Engineer
 ing\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
