236 lines
7.9 KiB
YAML
236 lines
7.9 KiB
YAML
venue: SC
|
||
year: 2025
|
||
papers:
|
||
- title: 'Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability'
|
||
authors:
|
||
- Nicholas Frontiere
|
||
- J. D. Emberson
|
||
- Michael Buehlmann
|
||
- Esteban M. Rangel
|
||
- Salman Habib 0002
|
||
- Katrin Heitmann
|
||
- Patricia Larsen
|
||
- Vitali A. Morozov
|
||
- Adrian Pope
|
||
- Claude-André Faucher-Giguère
|
||
- Antigoni Georgiadou
|
||
- Damien Lebrun-Grandié
|
||
- Andrey Prokopenko
|
||
reason: "Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application."
|
||
|
||
- title: Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance
|
||
authors:
|
||
- Nicolas Vetsch
|
||
- Alexander Maeder
|
||
- Vincent Maillou
|
||
- Anders Winka
|
||
- Jiang Cao
|
||
- Grzegorz Kwasniewski
|
||
- Leonard Deuschle
|
||
- Torsten Hoefler
|
||
- Alexandros Nikolaos Ziogas
|
||
- Mathieu Luisier
|
||
reason: "Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method."
|
||
|
||
- title: Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers
|
||
authors:
|
||
- Kai Xu
|
||
- Maoxue Yu
|
||
- Yuhu Chen
|
||
- Jie Gao
|
||
- Shuang Wang
|
||
- Jiaying Song
|
||
- Xiaohui Duan
|
||
- Junwei Wei
|
||
- Jiangfeng Yu
|
||
- Hailong Liu 0007
|
||
- Jinrong Jiang
|
||
- Yi Zhang 0127
|
||
- Pengfei Lin 0004
|
||
- Tianyi Wang
|
||
- Pengfei Wang
|
||
- Weipeng Zheng
|
||
- Jingwei Xie
|
||
- Jiakang Zhang
|
||
- Zilu Liu
|
||
- Xiaoyu Jin
|
||
- Jilin Wei
|
||
- Qixin Chang
|
||
- Qingxia Lin
|
||
- Yanzhi Zhou
|
||
- Weiguo Liu
|
||
- Wei Xue 0003
|
||
- Yiwen Li
|
||
- Haohuan Fu
|
||
- Yue Yu 0001
|
||
- Xuebin Chi
|
||
- Lixin Wu
|
||
reason: "Integrates AI acceleration into a kilometer-scale climate model to reach year-scale simulation throughput, showing how ML–physics hybrid approaches can redefine climate modeling at supercomputer scale."
|
||
|
||
- title: 'Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity'
|
||
authors:
|
||
- Tommaso Bonato
|
||
- Sepehr Abdous
|
||
- Abdul Kabbani
|
||
- Ahmad Ghalayini
|
||
- Nadeen Gebara
|
||
- Terry Lam
|
||
- Anup Agarwal
|
||
- Tiancheng Chen
|
||
- Zhuolong Yu
|
||
- Konstantin Taranov
|
||
- Mahmoud Elhaddad
|
||
- Daniele De Sensi
|
||
- Soudeh Ghorbani
|
||
- Torsten Hoefler
|
||
reason: "Proposes a unified congestion control and reliable transport architecture spanning intra- and inter-datacenter links, with strong throughput and latency results relevant to AI and HPC clusters."
|
||
|
||
- title: 'SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication'
|
||
authors:
|
||
- Mikhail Khalilov
|
||
- Siyuan Shen
|
||
- Marcin Chrapek
|
||
- Tiancheng Chen
|
||
- Kenji Nakano
|
||
- Nicola Mazzoletti
|
||
- Peter-Jan Gootzen
|
||
- Salvatore Di Girolamo
|
||
- Rami Nudelman
|
||
- Gil Bloch
|
||
- Jithin Jose
|
||
- Abdul Kabbani
|
||
- Sreevatsa Anantharamu
|
||
- Jie Zhang
|
||
- Konstantin Taranov
|
||
- Zhuolong Yu
|
||
- Scott Moe
|
||
- Mahmoud Elhaddad
|
||
- Torsten Hoefler
|
||
reason: "Introduces software-defined reliability for RDMA at global scale, decoupling reliability policies from hardware to dramatically improve fault tolerance and reconfigurability in large-scale HPC networks."
|
||
|
||
- title: 'Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality'
|
||
authors:
|
||
- Daniele De Sensi
|
||
- Saverio Pasqualoni
|
||
- Lorenzo Piarulli
|
||
- Tommaso Bonato
|
||
- Seydou Ba
|
||
- Matteo Turisini
|
||
- Jens Domke
|
||
- Torsten Hoefler
|
||
reason: "Presents bine tree topologies for MPI collective operations that exploit communication locality, yielding significant latency and bandwidth improvements over standard binomial trees on modern HPC interconnects."
|
||
|
||
- title: 'STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems'
|
||
authors:
|
||
- Chris Egersdoerfer
|
||
- Philip H. Carns
|
||
- Shane Snyder
|
||
- Robert Ross
|
||
- Dong Dai 0001
|
||
reason: "Demonstrates that an LLM-driven autonomous reasoning engine can tune parallel file system parameters as effectively as expert hand-tuning, opening a new direction for self-optimizing HPC storage."
|
||
|
||
- title: 'Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony Buffers'
|
||
authors:
|
||
- Jianqin Yan
|
||
- Shi Qiu 0012
|
||
- Yina Lv
|
||
- Yifan Hu
|
||
- Hao Chen
|
||
- Zhirong Shen
|
||
- Xin Yao
|
||
- Renhai Chen
|
||
- Jiwu Shu
|
||
- Gong Zhang 0001
|
||
- Yiming Zhang 0003
|
||
reason: "Redesigns the GPU direct storage I/O stack to eliminate staging buffers, achieving large bandwidth gains for GPU-to-SSD transfers critical to LLM training and scientific data workflows."
|
||
|
||
- title: Breaking the System Noise Barrier at Exascale
|
||
authors:
|
||
- Edgar A. León
|
||
- Joseph Glenski
|
||
- Mark J. Stock
|
||
- Kim H. McMahon
|
||
- William Loewe
|
||
- Clark Snyder
|
||
- Larry Kaplan
|
||
- Srinath Vadlamani
|
||
- Timothy I. Mattox
|
||
- Trent D'Hooge
|
||
- Brian Behlendorf
|
||
- Nathan Hanford
|
||
- Ramesh Pankajakshan
|
||
- Matthew L. Leininger
|
||
reason: "Provides a rigorous characterization and mitigation of OS and hardware noise at exascale, demonstrating measurable improvements in collective communication performance on a real production system."
|
||
|
||
- title: 'Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs'
|
||
authors:
|
||
- Shengkun Cui
|
||
- Archit Patke
|
||
- Hung Nguyen
|
||
- Aditya Ranjan
|
||
- Ziheng Chen 0006
|
||
- Phuong Cao
|
||
- Gregory H. Bauer
|
||
- Brett M. Bode
|
||
- Catello Di Martino
|
||
- Saurabh Jha
|
||
- Chandra Narayanaswami
|
||
- Daby Sow
|
||
- Zbigniew T. Kalbarczyk
|
||
- Ravishankar K. Iyer
|
||
reason: "Delivers the first detailed side-by-side hardware fault-injection study of H100 and A100 GPUs, revealing how architecture changes in Hopper alter error propagation and resilience for HPC and AI workloads."
|
||
|
||
- title: Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems
|
||
authors:
|
||
- Pengfei Yu 0002
|
||
- Jingjing Gu
|
||
- Hao Han
|
||
- Dazhong Shen
|
||
- Bao Wen
|
||
- Yang Liu 0390
|
||
reason: "Characterizes real-world failure modes of large-scale LLM training on HPC clusters and proposes targeted mitigation strategies, providing essential reliability insights for AI infrastructure operators."
|
||
|
||
- title: 'XaaS Containers: Performance-Portable Representation With Source and IR Containers'
|
||
authors:
|
||
- Marcin Copik
|
||
- Eiman Alnuaimi
|
||
- Alok Kamatar
|
||
- Valérie Hayot-Sasson
|
||
- Alberto Madonna
|
||
- Todd Gamblin
|
||
- Kyle Chard
|
||
- Ian T. Foster
|
||
- Torsten Hoefler
|
||
reason: "Proposes source- and IR-level HPC containers that enable performance portability across heterogeneous architectures without recompilation, addressing a key deployment challenge for reproducible HPC software."
|
||
|
||
- title: 'cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications'
|
||
authors:
|
||
- Xi Wang 0027
|
||
- Bin Ma
|
||
- Jongryool Kim
|
||
- Byungil Koh
|
||
- Hoshik Kim
|
||
- Dong Li 0001
|
||
reason: "Exploits CXL memory semantics to implement MPI communication primitives with dramatically reduced software overhead, demonstrating a promising path for memory-centric supercomputer interconnects."
|
||
|
||
- title: 'X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms'
|
||
authors:
|
||
- Yueming Yuan
|
||
- Ahan Gupta
|
||
- Jianping Li
|
||
- Sajal Dash
|
||
- Feiyi Wang
|
||
- Minjia Zhang
|
||
reason: "Addresses the communication and load-balance bottlenecks of sparse Mixture-of-Experts training at scale, achieving efficient utilization of large GPU clusters for next-generation LLM workloads."
|
||
|
||
- title: Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing
|
||
authors:
|
||
- Oscar Antepara
|
||
- Zhengji Zhao
|
||
- Brian Austin
|
||
- Nan Ding 0006
|
||
- Leonid Oliker
|
||
- Nicholas J. Wright
|
||
- Samuel Williams 0001
|
||
reason: "Develops fine-grained benchmark-driven energy models for GPU supercomputers that attribute power consumption to individual components and workloads, enabling principled energy optimization at the facility level."
|