venue: SC year: 2024 papers: - title: "Pushing the Limit of Quantum Mechanical Simulation to the Raman Spectra of\ \ a Biological System with 100 Million Atoms" authors: - Honghui Shang - Ying Liu 0055 - Zhikun Wu - Zhenchuan Chen - Jinfeng Liu 0004 - Meiyue Shao - Yingzhou Li - Bowen Kan - Huimin Cui - Xiaobing Feng 0002 - Yunquan Zhang - Donald G. Truhlar - Hong An - Xiao He 0004 - Jinlong Yang 0003 reason: "Gordon Bell-class result scaling ab initio Raman spectroscopy to 100 million atoms, pushing quantum-chemical simulation well beyond prior limits." - title: "Breaking the Molecular Dynamics Timescale Barrier Using a Wafer-Scale System" authors: - Kylee Santos - Stan G. Moore - Tomas Oppelstrup - Amirali Sharifian - Ilya Sharapov - Aidan P. Thompson - Delyan Z. Kalchev - Danny Perez - Robert Schreiber - Scott Pakin - Edgar A. Leon - James H. Laros III - Michael James 0002 - Sivasankaran Rajamanickam reason: "Demonstrates how a Cerebras wafer-scale engine shatters the classical MD timescale barrier, enabling microsecond-regime atomistic simulation at unprecedented speed." - title: "Scaling Molecular Dynamics with ab initio Accuracy to 149 Nanoseconds per\ \ Day" authors: - Jianxiong Li - Boyang Li - Zhuoqiang Guo - Mingzhen Li 0001 - Enji Li - Lijun Liu - Guojun Yuan - Zhan Wang 0003 - Guangming Tan - Weile Jia reason: "Achieves 149 ns/day for large-scale deep-potential MD, combining neural-network potentials and HPC engineering to approach DFT accuracy at AIMD-like scale." - title: "Breaking the Million-Electron and 1 EFLOP/s Barriers: Biomolecular-Scale\ \ Ab Initio Molecular Dynamics Using MP2 Potentials" authors: - Ryan Stocks - Jorge L. Galvez Vallejo - Fiona C. Y. Yu - Calum Snowdon - Elise Palethorpe - Jakub Kurzak - Dmytro Bykov - Giuseppe M. J. Barca reason: "First demonstration of MP2-level AIMD at the million-electron and exaFLOP/s scale, a landmark in quantum chemistry on supercomputers." - title: "Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep\ \ Learning" authors: - Wei An - Xiao Bi - Guanting Chen 0002 - Shanhuang Chen - Chengqi Deng - Honghui Ding - Kai Dong 0003 - Qiushi Du - Wenjun Gao - Kang Guan - Jianzhong Guo - Yongqiang Guo - Zhe Fu 0009 - Ying He 0018 - Panpan Huang - Jiashi Li - Wenfeng Liang - Xiaodong Liu 0021 - Xin Liu 0126 - Yiyuan Liu - Yuxuan Liu 0019 - Shanghao Lu - Xuan Lu - Xiaotao Nie - Tian Pei - Junjie Qiu - Hui Qu - Zehui Ren - Zhangli Sha - Xuecheng Su - Xiaowen Sun - Yixuan Tan - Minghui Tang - Shiyu Wang - Yaohui Wang - Yongji Wang - Ziwei Xie - Yiliang Xiong - Yanhong Xu - Shengfeng Ye - Shuiping Yu - Yukun Zha - Liyue Zhang - Haowei Zhang - Mingchuan Zhang - Wentao Zhang - Yichao Zhang 0004 - Chenggang Zhao - Yao Zhao 0005 - Shangyan Zhou - Shunfeng Zhou - Yuheng Zou reason: "Full system co-design report from DeepSeek's AI-HPC cluster showing 40% cost reduction vs. NVIDIA DGX through network/software optimizations, with production evidence at scale." - title: "MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows\ \ with Direct Preference Optimization" authors: - Gautham Dharuman - Kyle Hippe - Alexander Brace - Sam Foreman - Väinö Hatanpää - Varuni Katti Sastry - Huihuo Zheng - Logan T. Ward - Servesh Muralidharan - Archit Vasan - Bharat Kale - Carla M. Mann - Heng Ma - Yun-Hsuan Cheng - Yuliana Zamora - Shengchao Liu - Chaowei Xiao - Murali Emani - Tom Gibbs - Mahidhar Tatineni - Deepak Canchi - Jerome Mitchell - Koichi Yamada - Maria Garzaran 0001 - Michael E. Papka - Ian T. Foster - Rick Stevens - Anima Anandkumar - Venkatram Vishwanath - Arvind Ramanathan reason: "First exaFLOP/s AI science workflow, integrating multimodal protein design with DPO alignment at supercomputing scale across Frontier and Aurora." - title: "ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability" authors: - Xiao Wang 0004 - Siyan Liu - Aristeidis Tsaris - Jong-Youl Choi - Ashwin M. Aji - Ming Fan - Wei Zhang 0261 - Junqi Yin - Moetasim Ashfaq - Dan Lu 0001 - Prasanna Balaprakash reason: "Introduces a large foundation model for Earth system prediction trained on Frontier, demonstrating how exascale AI infrastructure enables climate-scale spatiotemporal modeling." - title: "Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers" authors: - Siddharth Singh - Prajwal Singhania - Aditya K. Ranjan - John Kirchenbauer - Jonas Geiping - Yuxin Wen - Neel Jain - Abhimanyu Hans - Manli Shu - Aditya Tomar - Tom Goldstein - Abhinav Bhatele reason: "Presents an open-source framework for LLM training at thousands-of-GPU scale, systematically analyzing throughput, memory, and communication trade-offs on leadership supercomputers." - title: "Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects" authors: - Daniele De Sensi - Lorenzo Pichetti - Flavio Vella - Tiziano De Matteis - Zebin Ren - Luigi Fusco - Matteo Turisini - Daniele Cesarini - Kurt Lust - Animesh Trivedi - Duncan Roweth - Filippo Spiga - Salvatore Di Girolamo - Torsten Hoefler reason: "Comprehensive empirical study of GPU-to-GPU communication across six major supercomputers, revealing bottlenecks and bandwidth characteristics relevant to all distributed AI/HPC workloads." - title: "Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed\ \ AI" authors: - Mikhail Khalilov - Salvatore Di Girolamo - Marcin Chrapek - Rami Nudelman - Gil Bloch - Torsten Hoefler reason: "Achieves bandwidth-optimal collective communication by offloading broadcast and allgather to SmartNICs, directly benefiting large-scale distributed deep learning." - title: "A Workflow Roofline Model for End-to-End Workflow Performance Analysis" authors: - Nan Ding 0006 - Brian Austin - Yang Liu 0179 - Neil Mehta - Steven Farrell - Johannes P. Blaschke - Leonid Oliker - Hai Ah Nam - Nicholas J. Wright - Samuel Williams 0001 reason: "Extends the Roofline model to full end-to-end HPC workflows, enabling systematic performance diagnosis across compute, I/O, and data movement stages." - title: "GVARP: Detecting Performance Variance on Large-Scale Heterogeneous Systems" authors: - Xin You 0001 - Zhibo Xuan - Hailong Yang 0002 - Zhongzhi Luan - Yi Liu 0013 - Depei Qian 0002 reason: "Identifies and diagnoses GPU performance variance at scale on heterogeneous supercomputers, an increasingly critical issue for reproducibility and efficiency." - title: "A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated\ \ at Exascale" authors: - Wesley Brewer - Matthias Maiterth - Vineet Kumar - Rafal P. Wojda - Sedrick Bouknight - Jesse Hines - Woong Shin - Scott Greenwood - David Grant - Wesley Williams - Feiyi Wang reason: "First deployment of a digital twin for a liquid-cooled exascale system (Frontier), enabling real-time thermal and power management with validated empirical results." - title: "Doubling Graph Traversal Efficiency to 198 TeraTEPS on the Supercomputer\ \ Fugaku" authors: - Junya Arai - Masahiro Nakao - Yuto Inoue - Kanto Teranishi - Koji Ueno - Keiichiro Yamamura - Mitsuhisa Sato - Katsuki Fujisawa reason: "Sets a new world record for graph traversal at 198 TTEPS on Fugaku through novel communication and load-balancing techniques, a landmark Graph500 result." - title: "MegaMmap: Blurring the Boundary Between Memory and Storage for Data-Intensive\ \ Workloads" authors: - Luke Logan - Anthony Kougkas - Xian-He Sun reason: "Novel storage abstraction that transparently tiered memory and storage hierarchies, delivering near-DRAM performance for data-intensive HPC and AI workloads."