BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T054803Z
LOCATION:225-226
DTSTART;TZID=America/Chicago:20211117T103000
DTEND;TZID=America/Chicago:20211117T110000
UID:submissions.supercomputing.org_SC21_sess151_pap357@linklings.com
SUMMARY:Hardware Acceleration of Tensor-Structured Multilevel Ewald Summat
 ion Method on MDGRAPE-4A, a Special-Purpose Computer System for Molecular 
 Dynamics Simulations
DESCRIPTION:Paper\n\nHardware Acceleration of Tensor-Structured Multilevel
  Ewald Summation Method on MDGRAPE-4A, a Special-Purpose Computer System f
 or Molecular Dynamics Simulations\n\nMorimoto, Koyama, Zhang, Komatsu, Ohn
 o...\n\nWe developed MDGRAPE-4A, a special-purpose computer system for mol
 ecular dynamics simulations, consisting of 512 nodes of custom system-on-a
 -chip LSIs with dedicated processor cores and interconnects designed to ac
 hieve strong scalability for biomolecular simulations. To reduce the globa
 l communications required for the evaluation of Coulomb interactions, we c
 onducted a co-design of the MDGRAPE-4A and the novel algorithm, tensor-str
 uctured multilevel Ewald summation method (TME), which produced hardware m
 odules on the custom LSI circuit for particle–grid operations and for grid
 –grid separable convolutions on a 3D torus network. We implemented the con
 volution for the top-level grid potentials by using 3D FFTs on an FPGA, al
 ong with an FPGA-based octree network to gather grid charges. The elapsed 
 time for the long-range part of Coulomb is 50 &#956;s, which can mostly ov
 erlap with those for the short-range part, and the additional cost is appr
 oximately 10 &#956;s/step, which is only a 5% performance loss.\n\nTag: Ac
 celerator-based Architectures\n\nRegistration Category: Tech Program Reg P
 ass
END:VEVENT
END:VCALENDAR
