venue: SC year: 2025 papers: - title: 'Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability' authors: - Nicholas Frontiere - J. D. Emberson - Michael Buehlmann - Esteban M. Rangel - Salman Habib 0002 - Katrin Heitmann - Patricia Larsen - Vitali A. Morozov - Adrian Pope - Claude-André Faucher-Giguère - Antigoni Georgiadou - Damien Lebrun-Grandié - Andrey Prokopenko reason: "Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application." - title: Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance authors: - Nicolas Vetsch - Alexander Maeder - Vincent Maillou - Anders Winka - Jiang Cao - Grzegorz Kwasniewski - Leonard Deuschle - Torsten Hoefler - Alexandros Nikolaos Ziogas - Mathieu Luisier reason: "Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method." - title: Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers authors: - Kai Xu - Maoxue Yu - Yuhu Chen - Jie Gao - Shuang Wang - Jiaying Song - Xiaohui Duan - Junwei Wei - Jiangfeng Yu - Hailong Liu 0007 - Jinrong Jiang - Yi Zhang 0127 - Pengfei Lin 0004 - Tianyi Wang - Pengfei Wang - Weipeng Zheng - Jingwei Xie - Jiakang Zhang - Zilu Liu - Xiaoyu Jin - Jilin Wei - Qixin Chang - Qingxia Lin - Yanzhi Zhou - Weiguo Liu - Wei Xue 0003 - Yiwen Li - Haohuan Fu - Yue Yu 0001 - Xuebin Chi - Lixin Wu reason: "Integrates AI acceleration into a kilometer-scale climate model to reach year-scale simulation throughput, showing how ML–physics hybrid approaches can redefine climate modeling at supercomputer scale." - title: 'Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity' authors: - Tommaso Bonato - Sepehr Abdous - Abdul Kabbani - Ahmad Ghalayini - Nadeen Gebara - Terry Lam - Anup Agarwal - Tiancheng Chen - Zhuolong Yu - Konstantin Taranov - Mahmoud Elhaddad - Daniele De Sensi - Soudeh Ghorbani - Torsten Hoefler reason: "Proposes a unified congestion control and reliable transport architecture spanning intra- and inter-datacenter links, with strong throughput and latency results relevant to AI and HPC clusters." - title: 'SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication' authors: - Mikhail Khalilov - Siyuan Shen - Marcin Chrapek - Tiancheng Chen - Kenji Nakano - Nicola Mazzoletti - Peter-Jan Gootzen - Salvatore Di Girolamo - Rami Nudelman - Gil Bloch - Jithin Jose - Abdul Kabbani - Sreevatsa Anantharamu - Jie Zhang - Konstantin Taranov - Zhuolong Yu - Scott Moe - Mahmoud Elhaddad - Torsten Hoefler reason: "Introduces software-defined reliability for RDMA at global scale, decoupling reliability policies from hardware to dramatically improve fault tolerance and reconfigurability in large-scale HPC networks." - title: 'Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality' authors: - Daniele De Sensi - Saverio Pasqualoni - Lorenzo Piarulli - Tommaso Bonato - Seydou Ba - Matteo Turisini - Jens Domke - Torsten Hoefler reason: "Presents bine tree topologies for MPI collective operations that exploit communication locality, yielding significant latency and bandwidth improvements over standard binomial trees on modern HPC interconnects." - title: 'STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems' authors: - Chris Egersdoerfer - Philip H. Carns - Shane Snyder - Robert Ross - Dong Dai 0001 reason: "Demonstrates that an LLM-driven autonomous reasoning engine can tune parallel file system parameters as effectively as expert hand-tuning, opening a new direction for self-optimizing HPC storage." - title: 'Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony Buffers' authors: - Jianqin Yan - Shi Qiu 0012 - Yina Lv - Yifan Hu - Hao Chen - Zhirong Shen - Xin Yao - Renhai Chen - Jiwu Shu - Gong Zhang 0001 - Yiming Zhang 0003 reason: "Redesigns the GPU direct storage I/O stack to eliminate staging buffers, achieving large bandwidth gains for GPU-to-SSD transfers critical to LLM training and scientific data workflows." - title: Breaking the System Noise Barrier at Exascale authors: - Edgar A. León - Joseph Glenski - Mark J. Stock - Kim H. McMahon - William Loewe - Clark Snyder - Larry Kaplan - Srinath Vadlamani - Timothy I. Mattox - Trent D'Hooge - Brian Behlendorf - Nathan Hanford - Ramesh Pankajakshan - Matthew L. Leininger reason: "Provides a rigorous characterization and mitigation of OS and hardware noise at exascale, demonstrating measurable improvements in collective communication performance on a real production system." - title: 'Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs' authors: - Shengkun Cui - Archit Patke - Hung Nguyen - Aditya Ranjan - Ziheng Chen 0006 - Phuong Cao - Gregory H. Bauer - Brett M. Bode - Catello Di Martino - Saurabh Jha - Chandra Narayanaswami - Daby Sow - Zbigniew T. Kalbarczyk - Ravishankar K. Iyer reason: "Delivers the first detailed side-by-side hardware fault-injection study of H100 and A100 GPUs, revealing how architecture changes in Hopper alter error propagation and resilience for HPC and AI workloads." - title: Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems authors: - Pengfei Yu 0002 - Jingjing Gu - Hao Han - Dazhong Shen - Bao Wen - Yang Liu 0390 reason: "Characterizes real-world failure modes of large-scale LLM training on HPC clusters and proposes targeted mitigation strategies, providing essential reliability insights for AI infrastructure operators." - title: 'XaaS Containers: Performance-Portable Representation With Source and IR Containers' authors: - Marcin Copik - Eiman Alnuaimi - Alok Kamatar - Valérie Hayot-Sasson - Alberto Madonna - Todd Gamblin - Kyle Chard - Ian T. Foster - Torsten Hoefler reason: "Proposes source- and IR-level HPC containers that enable performance portability across heterogeneous architectures without recompilation, addressing a key deployment challenge for reproducible HPC software." - title: 'cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications' authors: - Xi Wang 0027 - Bin Ma - Jongryool Kim - Byungil Koh - Hoshik Kim - Dong Li 0001 reason: "Exploits CXL memory semantics to implement MPI communication primitives with dramatically reduced software overhead, demonstrating a promising path for memory-centric supercomputer interconnects." - title: 'X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms' authors: - Yueming Yuan - Ahan Gupta - Jianping Li - Sajal Dash - Feiyi Wang - Minjia Zhang reason: "Addresses the communication and load-balance bottlenecks of sparse Mixture-of-Experts training at scale, achieving efficient utilization of large GPU clusters for next-generation LLM workloads." - title: Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing authors: - Oscar Antepara - Zhengji Zhao - Brian Austin - Nan Ding 0006 - Leonid Oliker - Nicholas J. Wright - Samuel Williams 0001 reason: "Develops fine-grained benchmark-driven energy models for GPU supercomputers that attribute power consumption to individual components and workloads, enabling principled energy optimization at the facility level."