Files
publish-assistant/site/data/cloud-edge/papers/SC-2025-digest.yaml
Vincent Lannurien d822cdaa6a
All checks were successful
Build and deploy static pages / build-and-push (push) Successful in 19s
multi-topic, publish from gh-pages branch
2026-08-17 18:10:49 +02:00

236 lines
7.9 KiB
YAML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
venue: SC
year: 2025
papers:
- title: 'Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability'
authors:
- Nicholas Frontiere
- J. D. Emberson
- Michael Buehlmann
- Esteban M. Rangel
- Salman Habib 0002
- Katrin Heitmann
- Patricia Larsen
- Vitali A. Morozov
- Adrian Pope
- Claude-André Faucher-Giguère
- Antigoni Georgiadou
- Damien Lebrun-Grandié
- Andrey Prokopenko
reason: "Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application."
- title: Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance
authors:
- Nicolas Vetsch
- Alexander Maeder
- Vincent Maillou
- Anders Winka
- Jiang Cao
- Grzegorz Kwasniewski
- Leonard Deuschle
- Torsten Hoefler
- Alexandros Nikolaos Ziogas
- Mathieu Luisier
reason: "Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method."
- title: Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers
authors:
- Kai Xu
- Maoxue Yu
- Yuhu Chen
- Jie Gao
- Shuang Wang
- Jiaying Song
- Xiaohui Duan
- Junwei Wei
- Jiangfeng Yu
- Hailong Liu 0007
- Jinrong Jiang
- Yi Zhang 0127
- Pengfei Lin 0004
- Tianyi Wang
- Pengfei Wang
- Weipeng Zheng
- Jingwei Xie
- Jiakang Zhang
- Zilu Liu
- Xiaoyu Jin
- Jilin Wei
- Qixin Chang
- Qingxia Lin
- Yanzhi Zhou
- Weiguo Liu
- Wei Xue 0003
- Yiwen Li
- Haohuan Fu
- Yue Yu 0001
- Xuebin Chi
- Lixin Wu
reason: "Integrates AI acceleration into a kilometer-scale climate model to reach year-scale simulation throughput, showing how MLphysics hybrid approaches can redefine climate modeling at supercomputer scale."
- title: 'Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity'
authors:
- Tommaso Bonato
- Sepehr Abdous
- Abdul Kabbani
- Ahmad Ghalayini
- Nadeen Gebara
- Terry Lam
- Anup Agarwal
- Tiancheng Chen
- Zhuolong Yu
- Konstantin Taranov
- Mahmoud Elhaddad
- Daniele De Sensi
- Soudeh Ghorbani
- Torsten Hoefler
reason: "Proposes a unified congestion control and reliable transport architecture spanning intra- and inter-datacenter links, with strong throughput and latency results relevant to AI and HPC clusters."
- title: 'SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication'
authors:
- Mikhail Khalilov
- Siyuan Shen
- Marcin Chrapek
- Tiancheng Chen
- Kenji Nakano
- Nicola Mazzoletti
- Peter-Jan Gootzen
- Salvatore Di Girolamo
- Rami Nudelman
- Gil Bloch
- Jithin Jose
- Abdul Kabbani
- Sreevatsa Anantharamu
- Jie Zhang
- Konstantin Taranov
- Zhuolong Yu
- Scott Moe
- Mahmoud Elhaddad
- Torsten Hoefler
reason: "Introduces software-defined reliability for RDMA at global scale, decoupling reliability policies from hardware to dramatically improve fault tolerance and reconfigurability in large-scale HPC networks."
- title: 'Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality'
authors:
- Daniele De Sensi
- Saverio Pasqualoni
- Lorenzo Piarulli
- Tommaso Bonato
- Seydou Ba
- Matteo Turisini
- Jens Domke
- Torsten Hoefler
reason: "Presents bine tree topologies for MPI collective operations that exploit communication locality, yielding significant latency and bandwidth improvements over standard binomial trees on modern HPC interconnects."
- title: 'STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems'
authors:
- Chris Egersdoerfer
- Philip H. Carns
- Shane Snyder
- Robert Ross
- Dong Dai 0001
reason: "Demonstrates that an LLM-driven autonomous reasoning engine can tune parallel file system parameters as effectively as expert hand-tuning, opening a new direction for self-optimizing HPC storage."
- title: 'Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony Buffers'
authors:
- Jianqin Yan
- Shi Qiu 0012
- Yina Lv
- Yifan Hu
- Hao Chen
- Zhirong Shen
- Xin Yao
- Renhai Chen
- Jiwu Shu
- Gong Zhang 0001
- Yiming Zhang 0003
reason: "Redesigns the GPU direct storage I/O stack to eliminate staging buffers, achieving large bandwidth gains for GPU-to-SSD transfers critical to LLM training and scientific data workflows."
- title: Breaking the System Noise Barrier at Exascale
authors:
- Edgar A. León
- Joseph Glenski
- Mark J. Stock
- Kim H. McMahon
- William Loewe
- Clark Snyder
- Larry Kaplan
- Srinath Vadlamani
- Timothy I. Mattox
- Trent D'Hooge
- Brian Behlendorf
- Nathan Hanford
- Ramesh Pankajakshan
- Matthew L. Leininger
reason: "Provides a rigorous characterization and mitigation of OS and hardware noise at exascale, demonstrating measurable improvements in collective communication performance on a real production system."
- title: 'Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs'
authors:
- Shengkun Cui
- Archit Patke
- Hung Nguyen
- Aditya Ranjan
- Ziheng Chen 0006
- Phuong Cao
- Gregory H. Bauer
- Brett M. Bode
- Catello Di Martino
- Saurabh Jha
- Chandra Narayanaswami
- Daby Sow
- Zbigniew T. Kalbarczyk
- Ravishankar K. Iyer
reason: "Delivers the first detailed side-by-side hardware fault-injection study of H100 and A100 GPUs, revealing how architecture changes in Hopper alter error propagation and resilience for HPC and AI workloads."
- title: Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems
authors:
- Pengfei Yu 0002
- Jingjing Gu
- Hao Han
- Dazhong Shen
- Bao Wen
- Yang Liu 0390
reason: "Characterizes real-world failure modes of large-scale LLM training on HPC clusters and proposes targeted mitigation strategies, providing essential reliability insights for AI infrastructure operators."
- title: 'XaaS Containers: Performance-Portable Representation With Source and IR Containers'
authors:
- Marcin Copik
- Eiman Alnuaimi
- Alok Kamatar
- Valérie Hayot-Sasson
- Alberto Madonna
- Todd Gamblin
- Kyle Chard
- Ian T. Foster
- Torsten Hoefler
reason: "Proposes source- and IR-level HPC containers that enable performance portability across heterogeneous architectures without recompilation, addressing a key deployment challenge for reproducible HPC software."
- title: 'cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications'
authors:
- Xi Wang 0027
- Bin Ma
- Jongryool Kim
- Byungil Koh
- Hoshik Kim
- Dong Li 0001
reason: "Exploits CXL memory semantics to implement MPI communication primitives with dramatically reduced software overhead, demonstrating a promising path for memory-centric supercomputer interconnects."
- title: 'X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms'
authors:
- Yueming Yuan
- Ahan Gupta
- Jianping Li
- Sajal Dash
- Feiyi Wang
- Minjia Zhang
reason: "Addresses the communication and load-balance bottlenecks of sparse Mixture-of-Experts training at scale, achieving efficient utilization of large GPU clusters for next-generation LLM workloads."
- title: Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing
authors:
- Oscar Antepara
- Zhengji Zhao
- Brian Austin
- Nan Ding 0006
- Leonid Oliker
- Nicholas J. Wright
- Samuel Williams 0001
reason: "Develops fine-grained benchmark-driven energy models for GPU supercomputers that attribute power consumption to individual components and workloads, enabling principled energy optimization at the facility level."