--- title: SC 2025 Digest venue: SC year: 2025 date: '2025-01-01' tags: [] paper_count: 15 draft: false --- 15 papers selected. --- ### Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability *Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban M. Rangel *et al.** **TL;DR** — Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application. --- ### Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance *Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka *et al.** **TL;DR** — Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method. --- ### Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers *Kai Xu, Maoxue Yu, Yuhu Chen, Jie Gao *et al.** **TL;DR** — Integrates AI acceleration into a kilometer-scale climate model to reach year-scale simulation throughput, showing how ML–physics hybrid approaches can redefine climate modeling at supercomputer scale. --- ### Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity *Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini *et al.** **TL;DR** — Proposes a unified congestion control and reliable transport architecture spanning intra- and inter-datacenter links, with strong throughput and latency results relevant to AI and HPC clusters. --- ### SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication *Mikhail Khalilov, Siyuan Shen, Marcin Chrapek, Tiancheng Chen *et al.** **TL;DR** — Introduces software-defined reliability for RDMA at global scale, decoupling reliability policies from hardware to dramatically improve fault tolerance and reconfigurability in large-scale HPC networks. --- ### Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality *Daniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli, Tommaso Bonato *et al.** **TL;DR** — Presents bine tree topologies for MPI collective operations that exploit communication locality, yielding significant latency and bandwidth improvements over standard binomial trees on modern HPC interconnects. --- ### STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems *Chris Egersdoerfer, Philip H. Carns, Shane Snyder, Robert Ross *et al.** **TL;DR** — Demonstrates that an LLM-driven autonomous reasoning engine can tune parallel file system parameters as effectively as expert hand-tuning, opening a new direction for self-optimizing HPC storage. --- ### Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony Buffers *Jianqin Yan, Shi Qiu 0012, Yina Lv, Yifan Hu *et al.** **TL;DR** — Redesigns the GPU direct storage I/O stack to eliminate staging buffers, achieving large bandwidth gains for GPU-to-SSD transfers critical to LLM training and scientific data workflows. --- ### Breaking the System Noise Barrier at Exascale *Edgar A. León, Joseph Glenski, Mark J. Stock, Kim H. McMahon *et al.** **TL;DR** — Provides a rigorous characterization and mitigation of OS and hardware noise at exascale, demonstrating measurable improvements in collective communication performance on a real production system. --- ### Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs *Shengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan *et al.** **TL;DR** — Delivers the first detailed side-by-side hardware fault-injection study of H100 and A100 GPUs, revealing how architecture changes in Hopper alter error propagation and resilience for HPC and AI workloads. --- ### Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems *Pengfei Yu 0002, Jingjing Gu, Hao Han, Dazhong Shen *et al.** **TL;DR** — Characterizes real-world failure modes of large-scale LLM training on HPC clusters and proposes targeted mitigation strategies, providing essential reliability insights for AI infrastructure operators. --- ### XaaS Containers: Performance-Portable Representation With Source and IR Containers *Marcin Copik, Eiman Alnuaimi, Alok Kamatar, Valérie Hayot-Sasson *et al.** **TL;DR** — Proposes source- and IR-level HPC containers that enable performance portability across heterogeneous architectures without recompilation, addressing a key deployment challenge for reproducible HPC software. --- ### cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications *Xi Wang 0027, Bin Ma, Jongryool Kim, Byungil Koh *et al.** **TL;DR** — Exploits CXL memory semantics to implement MPI communication primitives with dramatically reduced software overhead, demonstrating a promising path for memory-centric supercomputer interconnects. --- ### X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms *Yueming Yuan, Ahan Gupta, Jianping Li, Sajal Dash *et al.** **TL;DR** — Addresses the communication and load-balance bottlenecks of sparse Mixture-of-Experts training at scale, achieving efficient utilization of large GPU clusters for next-generation LLM workloads. --- ### Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing *Oscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding 0006 *et al.** **TL;DR** — Develops fine-grained benchmark-driven energy models for GPU supercomputers that attribute power consumption to individual components and workloads, enabling principled energy optimization at the facility level.