multi-topic, publish from gh-pages branch
All checks were successful
Build and deploy static pages / build-and-push (push) Successful in 19s
All checks were successful
Build and deploy static pages / build-and-push (push) Successful in 19s
This commit is contained in:
132
site/content/cloud-edge/digests/SC-2025/index.md
Normal file
132
site/content/cloud-edge/digests/SC-2025/index.md
Normal file
@@ -0,0 +1,132 @@
|
||||
---
|
||||
title: SC 2025 Digest
|
||||
venue: SC
|
||||
year: 2025
|
||||
date: '2025-01-01'
|
||||
tags: []
|
||||
paper_count: 15
|
||||
draft: false
|
||||
---
|
||||
|
||||
15 papers selected.
|
||||
|
||||
---
|
||||
|
||||
### Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability
|
||||
|
||||
*Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban M. Rangel *et al.**
|
||||
|
||||
**TL;DR** — Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application.
|
||||
|
||||
---
|
||||
|
||||
### Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance
|
||||
|
||||
*Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka *et al.**
|
||||
|
||||
**TL;DR** — Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method.
|
||||
|
||||
---
|
||||
|
||||
### Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers
|
||||
|
||||
*Kai Xu, Maoxue Yu, Yuhu Chen, Jie Gao *et al.**
|
||||
|
||||
**TL;DR** — Integrates AI acceleration into a kilometer-scale climate model to reach year-scale simulation throughput, showing how ML–physics hybrid approaches can redefine climate modeling at supercomputer scale.
|
||||
|
||||
---
|
||||
|
||||
### Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity
|
||||
|
||||
*Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini *et al.**
|
||||
|
||||
**TL;DR** — Proposes a unified congestion control and reliable transport architecture spanning intra- and inter-datacenter links, with strong throughput and latency results relevant to AI and HPC clusters.
|
||||
|
||||
---
|
||||
|
||||
### SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication
|
||||
|
||||
*Mikhail Khalilov, Siyuan Shen, Marcin Chrapek, Tiancheng Chen *et al.**
|
||||
|
||||
**TL;DR** — Introduces software-defined reliability for RDMA at global scale, decoupling reliability policies from hardware to dramatically improve fault tolerance and reconfigurability in large-scale HPC networks.
|
||||
|
||||
---
|
||||
|
||||
### Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
|
||||
|
||||
*Daniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli, Tommaso Bonato *et al.**
|
||||
|
||||
**TL;DR** — Presents bine tree topologies for MPI collective operations that exploit communication locality, yielding significant latency and bandwidth improvements over standard binomial trees on modern HPC interconnects.
|
||||
|
||||
---
|
||||
|
||||
### STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems
|
||||
|
||||
*Chris Egersdoerfer, Philip H. Carns, Shane Snyder, Robert Ross *et al.**
|
||||
|
||||
**TL;DR** — Demonstrates that an LLM-driven autonomous reasoning engine can tune parallel file system parameters as effectively as expert hand-tuning, opening a new direction for self-optimizing HPC storage.
|
||||
|
||||
---
|
||||
|
||||
### Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony Buffers
|
||||
|
||||
*Jianqin Yan, Shi Qiu 0012, Yina Lv, Yifan Hu *et al.**
|
||||
|
||||
**TL;DR** — Redesigns the GPU direct storage I/O stack to eliminate staging buffers, achieving large bandwidth gains for GPU-to-SSD transfers critical to LLM training and scientific data workflows.
|
||||
|
||||
---
|
||||
|
||||
### Breaking the System Noise Barrier at Exascale
|
||||
|
||||
*Edgar A. León, Joseph Glenski, Mark J. Stock, Kim H. McMahon *et al.**
|
||||
|
||||
**TL;DR** — Provides a rigorous characterization and mitigation of OS and hardware noise at exascale, demonstrating measurable improvements in collective communication performance on a real production system.
|
||||
|
||||
---
|
||||
|
||||
### Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
|
||||
|
||||
*Shengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan *et al.**
|
||||
|
||||
**TL;DR** — Delivers the first detailed side-by-side hardware fault-injection study of H100 and A100 GPUs, revealing how architecture changes in Hopper alter error propagation and resilience for HPC and AI workloads.
|
||||
|
||||
---
|
||||
|
||||
### Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems
|
||||
|
||||
*Pengfei Yu 0002, Jingjing Gu, Hao Han, Dazhong Shen *et al.**
|
||||
|
||||
**TL;DR** — Characterizes real-world failure modes of large-scale LLM training on HPC clusters and proposes targeted mitigation strategies, providing essential reliability insights for AI infrastructure operators.
|
||||
|
||||
---
|
||||
|
||||
### XaaS Containers: Performance-Portable Representation With Source and IR Containers
|
||||
|
||||
*Marcin Copik, Eiman Alnuaimi, Alok Kamatar, Valérie Hayot-Sasson *et al.**
|
||||
|
||||
**TL;DR** — Proposes source- and IR-level HPC containers that enable performance portability across heterogeneous architectures without recompilation, addressing a key deployment challenge for reproducible HPC software.
|
||||
|
||||
---
|
||||
|
||||
### cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
|
||||
|
||||
*Xi Wang 0027, Bin Ma, Jongryool Kim, Byungil Koh *et al.**
|
||||
|
||||
**TL;DR** — Exploits CXL memory semantics to implement MPI communication primitives with dramatically reduced software overhead, demonstrating a promising path for memory-centric supercomputer interconnects.
|
||||
|
||||
---
|
||||
|
||||
### X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
|
||||
|
||||
*Yueming Yuan, Ahan Gupta, Jianping Li, Sajal Dash *et al.**
|
||||
|
||||
**TL;DR** — Addresses the communication and load-balance bottlenecks of sparse Mixture-of-Experts training at scale, achieving efficient utilization of large GPU clusters for next-generation LLM workloads.
|
||||
|
||||
---
|
||||
|
||||
### Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing
|
||||
|
||||
*Oscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding 0006 *et al.**
|
||||
|
||||
**TL;DR** — Develops fine-grained benchmark-driven energy models for GPU supercomputers that attribute power consumption to individual components and workloads, enabling principled energy optimization at the facility level.
|
||||
|
||||
Reference in New Issue
Block a user