Files
publish-assistant/site/content/cloud-edge/digests/SC-2025/index.md
Vincent Lannurien d822cdaa6a
All checks were successful
Build and deploy static pages / build-and-push (push) Successful in 19s
multi-topic, publish from gh-pages branch
2026-08-17 18:10:49 +02:00

5.8 KiB
Raw Blame History

title, venue, year, date, tags, paper_count, draft
title venue year date tags paper_count draft
SC 2025 Digest SC 2025 2025-01-01
15 false

15 papers selected.


Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability

Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban M. Rangel et al.

TL;DR — Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application.


Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance

Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka et al.

TL;DR — Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method.


Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers

Kai Xu, Maoxue Yu, Yuhu Chen, Jie Gao et al.

TL;DR — Integrates AI acceleration into a kilometer-scale climate model to reach year-scale simulation throughput, showing how MLphysics hybrid approaches can redefine climate modeling at supercomputer scale.


Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity

Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini et al.

TL;DR — Proposes a unified congestion control and reliable transport architecture spanning intra- and inter-datacenter links, with strong throughput and latency results relevant to AI and HPC clusters.


SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication

Mikhail Khalilov, Siyuan Shen, Marcin Chrapek, Tiancheng Chen et al.

TL;DR — Introduces software-defined reliability for RDMA at global scale, decoupling reliability policies from hardware to dramatically improve fault tolerance and reconfigurability in large-scale HPC networks.


Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality

Daniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli, Tommaso Bonato et al.

TL;DR — Presents bine tree topologies for MPI collective operations that exploit communication locality, yielding significant latency and bandwidth improvements over standard binomial trees on modern HPC interconnects.


STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems

Chris Egersdoerfer, Philip H. Carns, Shane Snyder, Robert Ross et al.

TL;DR — Demonstrates that an LLM-driven autonomous reasoning engine can tune parallel file system parameters as effectively as expert hand-tuning, opening a new direction for self-optimizing HPC storage.


Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony Buffers

Jianqin Yan, Shi Qiu 0012, Yina Lv, Yifan Hu et al.

TL;DR — Redesigns the GPU direct storage I/O stack to eliminate staging buffers, achieving large bandwidth gains for GPU-to-SSD transfers critical to LLM training and scientific data workflows.


Breaking the System Noise Barrier at Exascale

Edgar A. León, Joseph Glenski, Mark J. Stock, Kim H. McMahon et al.

TL;DR — Provides a rigorous characterization and mitigation of OS and hardware noise at exascale, demonstrating measurable improvements in collective communication performance on a real production system.


Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs

Shengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan et al.

TL;DR — Delivers the first detailed side-by-side hardware fault-injection study of H100 and A100 GPUs, revealing how architecture changes in Hopper alter error propagation and resilience for HPC and AI workloads.


Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems

Pengfei Yu 0002, Jingjing Gu, Hao Han, Dazhong Shen et al.

TL;DR — Characterizes real-world failure modes of large-scale LLM training on HPC clusters and proposes targeted mitigation strategies, providing essential reliability insights for AI infrastructure operators.


XaaS Containers: Performance-Portable Representation With Source and IR Containers

Marcin Copik, Eiman Alnuaimi, Alok Kamatar, Valérie Hayot-Sasson et al.

TL;DR — Proposes source- and IR-level HPC containers that enable performance portability across heterogeneous architectures without recompilation, addressing a key deployment challenge for reproducible HPC software.


cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications

Xi Wang 0027, Bin Ma, Jongryool Kim, Byungil Koh et al.

TL;DR — Exploits CXL memory semantics to implement MPI communication primitives with dramatically reduced software overhead, demonstrating a promising path for memory-centric supercomputer interconnects.


X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms

Yueming Yuan, Ahan Gupta, Jianping Li, Sajal Dash et al.

TL;DR — Addresses the communication and load-balance bottlenecks of sparse Mixture-of-Experts training at scale, achieving efficient utilization of large GPU clusters for next-generation LLM workloads.


Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing

Oscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding 0006 et al.

TL;DR — Develops fine-grained benchmark-driven energy models for GPU supercomputers that attribute power consumption to individual components and workloads, enabling principled energy optimization at the facility level.