Files
Vincent Lannurien d822cdaa6a
All checks were successful
Build and deploy static pages / build-and-push (push) Successful in 19s
multi-topic, publish from gh-pages branch
2026-08-17 18:10:49 +02:00

133 lines
5.8 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
title: SC 2025 Digest
venue: SC
year: 2025
date: '2025-01-01'
tags: []
paper_count: 15
draft: false
---
15 papers selected.
---
### Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability
*Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban M. Rangel *et al.**
**TL;DR** — Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application.
---
### Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance
*Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka *et al.**
**TL;DR** — Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method.
---
### Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers
*Kai Xu, Maoxue Yu, Yuhu Chen, Jie Gao *et al.**
**TL;DR** — Integrates AI acceleration into a kilometer-scale climate model to reach year-scale simulation throughput, showing how MLphysics hybrid approaches can redefine climate modeling at supercomputer scale.
---
### Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity
*Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini *et al.**
**TL;DR** — Proposes a unified congestion control and reliable transport architecture spanning intra- and inter-datacenter links, with strong throughput and latency results relevant to AI and HPC clusters.
---
### SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication
*Mikhail Khalilov, Siyuan Shen, Marcin Chrapek, Tiancheng Chen *et al.**
**TL;DR** — Introduces software-defined reliability for RDMA at global scale, decoupling reliability policies from hardware to dramatically improve fault tolerance and reconfigurability in large-scale HPC networks.
---
### Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
*Daniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli, Tommaso Bonato *et al.**
**TL;DR** — Presents bine tree topologies for MPI collective operations that exploit communication locality, yielding significant latency and bandwidth improvements over standard binomial trees on modern HPC interconnects.
---
### STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems
*Chris Egersdoerfer, Philip H. Carns, Shane Snyder, Robert Ross *et al.**
**TL;DR** — Demonstrates that an LLM-driven autonomous reasoning engine can tune parallel file system parameters as effectively as expert hand-tuning, opening a new direction for self-optimizing HPC storage.
---
### Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony Buffers
*Jianqin Yan, Shi Qiu 0012, Yina Lv, Yifan Hu *et al.**
**TL;DR** — Redesigns the GPU direct storage I/O stack to eliminate staging buffers, achieving large bandwidth gains for GPU-to-SSD transfers critical to LLM training and scientific data workflows.
---
### Breaking the System Noise Barrier at Exascale
*Edgar A. León, Joseph Glenski, Mark J. Stock, Kim H. McMahon *et al.**
**TL;DR** — Provides a rigorous characterization and mitigation of OS and hardware noise at exascale, demonstrating measurable improvements in collective communication performance on a real production system.
---
### Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
*Shengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan *et al.**
**TL;DR** — Delivers the first detailed side-by-side hardware fault-injection study of H100 and A100 GPUs, revealing how architecture changes in Hopper alter error propagation and resilience for HPC and AI workloads.
---
### Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems
*Pengfei Yu 0002, Jingjing Gu, Hao Han, Dazhong Shen *et al.**
**TL;DR** — Characterizes real-world failure modes of large-scale LLM training on HPC clusters and proposes targeted mitigation strategies, providing essential reliability insights for AI infrastructure operators.
---
### XaaS Containers: Performance-Portable Representation With Source and IR Containers
*Marcin Copik, Eiman Alnuaimi, Alok Kamatar, Valérie Hayot-Sasson *et al.**
**TL;DR** — Proposes source- and IR-level HPC containers that enable performance portability across heterogeneous architectures without recompilation, addressing a key deployment challenge for reproducible HPC software.
---
### cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
*Xi Wang 0027, Bin Ma, Jongryool Kim, Byungil Koh *et al.**
**TL;DR** — Exploits CXL memory semantics to implement MPI communication primitives with dramatically reduced software overhead, demonstrating a promising path for memory-centric supercomputer interconnects.
---
### X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
*Yueming Yuan, Ahan Gupta, Jianping Li, Sajal Dash *et al.**
**TL;DR** — Addresses the communication and load-balance bottlenecks of sparse Mixture-of-Experts training at scale, achieving efficient utilization of large GPU clusters for next-generation LLM workloads.
---
### Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing
*Oscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding 0006 *et al.**
**TL;DR** — Develops fine-grained benchmark-driven energy models for GPU supercomputers that attribute power consumption to individual components and workloads, enabling principled energy optimization at the facility level.