multi-topic, publish from gh-pages branch
All checks were successful
Build and deploy static pages / build-and-push (push) Successful in 19s

This commit is contained in:
2026-08-17 18:10:49 +02:00
parent 1a9f822b56
commit d822cdaa6a
181 changed files with 1076 additions and 437 deletions

View File

@@ -0,0 +1,235 @@
venue: SC
year: 2025
papers:
- title: 'Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability'
authors:
- Nicholas Frontiere
- J. D. Emberson
- Michael Buehlmann
- Esteban M. Rangel
- Salman Habib 0002
- Katrin Heitmann
- Patricia Larsen
- Vitali A. Morozov
- Adrian Pope
- Claude-André Faucher-Giguère
- Antigoni Georgiadou
- Damien Lebrun-Grandié
- Andrey Prokopenko
reason: "Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application."
- title: Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance
authors:
- Nicolas Vetsch
- Alexander Maeder
- Vincent Maillou
- Anders Winka
- Jiang Cao
- Grzegorz Kwasniewski
- Leonard Deuschle
- Torsten Hoefler
- Alexandros Nikolaos Ziogas
- Mathieu Luisier
reason: "Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method."
- title: Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers
authors:
- Kai Xu
- Maoxue Yu
- Yuhu Chen
- Jie Gao
- Shuang Wang
- Jiaying Song
- Xiaohui Duan
- Junwei Wei
- Jiangfeng Yu
- Hailong Liu 0007
- Jinrong Jiang
- Yi Zhang 0127
- Pengfei Lin 0004
- Tianyi Wang
- Pengfei Wang
- Weipeng Zheng
- Jingwei Xie
- Jiakang Zhang
- Zilu Liu
- Xiaoyu Jin
- Jilin Wei
- Qixin Chang
- Qingxia Lin
- Yanzhi Zhou
- Weiguo Liu
- Wei Xue 0003
- Yiwen Li
- Haohuan Fu
- Yue Yu 0001
- Xuebin Chi
- Lixin Wu
reason: "Integrates AI acceleration into a kilometer-scale climate model to reach year-scale simulation throughput, showing how MLphysics hybrid approaches can redefine climate modeling at supercomputer scale."
- title: 'Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity'
authors:
- Tommaso Bonato
- Sepehr Abdous
- Abdul Kabbani
- Ahmad Ghalayini
- Nadeen Gebara
- Terry Lam
- Anup Agarwal
- Tiancheng Chen
- Zhuolong Yu
- Konstantin Taranov
- Mahmoud Elhaddad
- Daniele De Sensi
- Soudeh Ghorbani
- Torsten Hoefler
reason: "Proposes a unified congestion control and reliable transport architecture spanning intra- and inter-datacenter links, with strong throughput and latency results relevant to AI and HPC clusters."
- title: 'SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication'
authors:
- Mikhail Khalilov
- Siyuan Shen
- Marcin Chrapek
- Tiancheng Chen
- Kenji Nakano
- Nicola Mazzoletti
- Peter-Jan Gootzen
- Salvatore Di Girolamo
- Rami Nudelman
- Gil Bloch
- Jithin Jose
- Abdul Kabbani
- Sreevatsa Anantharamu
- Jie Zhang
- Konstantin Taranov
- Zhuolong Yu
- Scott Moe
- Mahmoud Elhaddad
- Torsten Hoefler
reason: "Introduces software-defined reliability for RDMA at global scale, decoupling reliability policies from hardware to dramatically improve fault tolerance and reconfigurability in large-scale HPC networks."
- title: 'Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality'
authors:
- Daniele De Sensi
- Saverio Pasqualoni
- Lorenzo Piarulli
- Tommaso Bonato
- Seydou Ba
- Matteo Turisini
- Jens Domke
- Torsten Hoefler
reason: "Presents bine tree topologies for MPI collective operations that exploit communication locality, yielding significant latency and bandwidth improvements over standard binomial trees on modern HPC interconnects."
- title: 'STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems'
authors:
- Chris Egersdoerfer
- Philip H. Carns
- Shane Snyder
- Robert Ross
- Dong Dai 0001
reason: "Demonstrates that an LLM-driven autonomous reasoning engine can tune parallel file system parameters as effectively as expert hand-tuning, opening a new direction for self-optimizing HPC storage."
- title: 'Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony Buffers'
authors:
- Jianqin Yan
- Shi Qiu 0012
- Yina Lv
- Yifan Hu
- Hao Chen
- Zhirong Shen
- Xin Yao
- Renhai Chen
- Jiwu Shu
- Gong Zhang 0001
- Yiming Zhang 0003
reason: "Redesigns the GPU direct storage I/O stack to eliminate staging buffers, achieving large bandwidth gains for GPU-to-SSD transfers critical to LLM training and scientific data workflows."
- title: Breaking the System Noise Barrier at Exascale
authors:
- Edgar A. León
- Joseph Glenski
- Mark J. Stock
- Kim H. McMahon
- William Loewe
- Clark Snyder
- Larry Kaplan
- Srinath Vadlamani
- Timothy I. Mattox
- Trent D'Hooge
- Brian Behlendorf
- Nathan Hanford
- Ramesh Pankajakshan
- Matthew L. Leininger
reason: "Provides a rigorous characterization and mitigation of OS and hardware noise at exascale, demonstrating measurable improvements in collective communication performance on a real production system."
- title: 'Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs'
authors:
- Shengkun Cui
- Archit Patke
- Hung Nguyen
- Aditya Ranjan
- Ziheng Chen 0006
- Phuong Cao
- Gregory H. Bauer
- Brett M. Bode
- Catello Di Martino
- Saurabh Jha
- Chandra Narayanaswami
- Daby Sow
- Zbigniew T. Kalbarczyk
- Ravishankar K. Iyer
reason: "Delivers the first detailed side-by-side hardware fault-injection study of H100 and A100 GPUs, revealing how architecture changes in Hopper alter error propagation and resilience for HPC and AI workloads."
- title: Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems
authors:
- Pengfei Yu 0002
- Jingjing Gu
- Hao Han
- Dazhong Shen
- Bao Wen
- Yang Liu 0390
reason: "Characterizes real-world failure modes of large-scale LLM training on HPC clusters and proposes targeted mitigation strategies, providing essential reliability insights for AI infrastructure operators."
- title: 'XaaS Containers: Performance-Portable Representation With Source and IR Containers'
authors:
- Marcin Copik
- Eiman Alnuaimi
- Alok Kamatar
- Valérie Hayot-Sasson
- Alberto Madonna
- Todd Gamblin
- Kyle Chard
- Ian T. Foster
- Torsten Hoefler
reason: "Proposes source- and IR-level HPC containers that enable performance portability across heterogeneous architectures without recompilation, addressing a key deployment challenge for reproducible HPC software."
- title: 'cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications'
authors:
- Xi Wang 0027
- Bin Ma
- Jongryool Kim
- Byungil Koh
- Hoshik Kim
- Dong Li 0001
reason: "Exploits CXL memory semantics to implement MPI communication primitives with dramatically reduced software overhead, demonstrating a promising path for memory-centric supercomputer interconnects."
- title: 'X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms'
authors:
- Yueming Yuan
- Ahan Gupta
- Jianping Li
- Sajal Dash
- Feiyi Wang
- Minjia Zhang
reason: "Addresses the communication and load-balance bottlenecks of sparse Mixture-of-Experts training at scale, achieving efficient utilization of large GPU clusters for next-generation LLM workloads."
- title: Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing
authors:
- Oscar Antepara
- Zhengji Zhao
- Brian Austin
- Nan Ding 0006
- Leonid Oliker
- Nicholas J. Wright
- Samuel Williams 0001
reason: "Develops fine-grained benchmark-driven energy models for GPU supercomputers that attribute power consumption to individual components and workloads, enabling principled energy optimization at the facility level."