content updates, various fixes

This commit is contained in:
khannurien
2026-04-26 12:57:40 +00:00
parent 8484abea47
commit 1a9f822b56
164 changed files with 82726 additions and 163 deletions

View File

@@ -0,0 +1,281 @@
venue: SC
year: 2024
papers:
- title: "Pushing the Limit of Quantum Mechanical Simulation to the Raman Spectra of\
\ a Biological System with 100 Million Atoms"
authors:
- Honghui Shang
- Ying Liu 0055
- Zhikun Wu
- Zhenchuan Chen
- Jinfeng Liu 0004
- Meiyue Shao
- Yingzhou Li
- Bowen Kan
- Huimin Cui
- Xiaobing Feng 0002
- Yunquan Zhang
- Donald G. Truhlar
- Hong An
- Xiao He 0004
- Jinlong Yang 0003
reason: "Gordon Bell-class result scaling ab initio Raman spectroscopy to 100 million atoms, pushing quantum-chemical simulation well beyond prior limits."
- title: "Breaking the Molecular Dynamics Timescale Barrier Using a Wafer-Scale System"
authors:
- Kylee Santos
- Stan G. Moore
- Tomas Oppelstrup
- Amirali Sharifian
- Ilya Sharapov
- Aidan P. Thompson
- Delyan Z. Kalchev
- Danny Perez
- Robert Schreiber
- Scott Pakin
- Edgar A. Leon
- James H. Laros III
- Michael James 0002
- Sivasankaran Rajamanickam
reason: "Demonstrates how a Cerebras wafer-scale engine shatters the classical MD timescale barrier, enabling microsecond-regime atomistic simulation at unprecedented speed."
- title: "Scaling Molecular Dynamics with ab initio Accuracy to 149 Nanoseconds per\
\ Day"
authors:
- Jianxiong Li
- Boyang Li
- Zhuoqiang Guo
- Mingzhen Li 0001
- Enji Li
- Lijun Liu
- Guojun Yuan
- Zhan Wang 0003
- Guangming Tan
- Weile Jia
reason: "Achieves 149 ns/day for large-scale deep-potential MD, combining neural-network potentials and HPC engineering to approach DFT accuracy at AIMD-like scale."
- title: "Breaking the Million-Electron and 1 EFLOP/s Barriers: Biomolecular-Scale\
\ Ab Initio Molecular Dynamics Using MP2 Potentials"
authors:
- Ryan Stocks
- Jorge L. Galvez Vallejo
- Fiona C. Y. Yu
- Calum Snowdon
- Elise Palethorpe
- Jakub Kurzak
- Dmytro Bykov
- Giuseppe M. J. Barca
reason: "First demonstration of MP2-level AIMD at the million-electron and exaFLOP/s scale, a landmark in quantum chemistry on supercomputers."
- title: "Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep\
\ Learning"
authors:
- Wei An
- Xiao Bi
- Guanting Chen 0002
- Shanhuang Chen
- Chengqi Deng
- Honghui Ding
- Kai Dong 0003
- Qiushi Du
- Wenjun Gao
- Kang Guan
- Jianzhong Guo
- Yongqiang Guo
- Zhe Fu 0009
- Ying He 0018
- Panpan Huang
- Jiashi Li
- Wenfeng Liang
- Xiaodong Liu 0021
- Xin Liu 0126
- Yiyuan Liu
- Yuxuan Liu 0019
- Shanghao Lu
- Xuan Lu
- Xiaotao Nie
- Tian Pei
- Junjie Qiu
- Hui Qu
- Zehui Ren
- Zhangli Sha
- Xuecheng Su
- Xiaowen Sun
- Yixuan Tan
- Minghui Tang
- Shiyu Wang
- Yaohui Wang
- Yongji Wang
- Ziwei Xie
- Yiliang Xiong
- Yanhong Xu
- Shengfeng Ye
- Shuiping Yu
- Yukun Zha
- Liyue Zhang
- Haowei Zhang
- Mingchuan Zhang
- Wentao Zhang
- Yichao Zhang 0004
- Chenggang Zhao
- Yao Zhao 0005
- Shangyan Zhou
- Shunfeng Zhou
- Yuheng Zou
reason: "Full system co-design report from DeepSeek's AI-HPC cluster showing 40% cost reduction vs. NVIDIA DGX through network/software optimizations, with production evidence at scale."
- title: "MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows\
\ with Direct Preference Optimization"
authors:
- Gautham Dharuman
- Kyle Hippe
- Alexander Brace
- Sam Foreman
- Väinö Hatanpää
- Varuni Katti Sastry
- Huihuo Zheng
- Logan T. Ward
- Servesh Muralidharan
- Archit Vasan
- Bharat Kale
- Carla M. Mann
- Heng Ma
- Yun-Hsuan Cheng
- Yuliana Zamora
- Shengchao Liu
- Chaowei Xiao
- Murali Emani
- Tom Gibbs
- Mahidhar Tatineni
- Deepak Canchi
- Jerome Mitchell
- Koichi Yamada
- Maria Garzaran 0001
- Michael E. Papka
- Ian T. Foster
- Rick Stevens
- Anima Anandkumar
- Venkatram Vishwanath
- Arvind Ramanathan
reason: "First exaFLOP/s AI science workflow, integrating multimodal protein design with DPO alignment at supercomputing scale across Frontier and Aurora."
- title: "ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability"
authors:
- Xiao Wang 0004
- Siyan Liu
- Aristeidis Tsaris
- Jong-Youl Choi
- Ashwin M. Aji
- Ming Fan
- Wei Zhang 0261
- Junqi Yin
- Moetasim Ashfaq
- Dan Lu 0001
- Prasanna Balaprakash
reason: "Introduces a large foundation model for Earth system prediction trained on Frontier, demonstrating how exascale AI infrastructure enables climate-scale spatiotemporal modeling."
- title: "Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers"
authors:
- Siddharth Singh
- Prajwal Singhania
- Aditya K. Ranjan
- John Kirchenbauer
- Jonas Geiping
- Yuxin Wen
- Neel Jain
- Abhimanyu Hans
- Manli Shu
- Aditya Tomar
- Tom Goldstein
- Abhinav Bhatele
reason: "Presents an open-source framework for LLM training at thousands-of-GPU scale, systematically analyzing throughput, memory, and communication trade-offs on leadership supercomputers."
- title: "Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects"
authors:
- Daniele De Sensi
- Lorenzo Pichetti
- Flavio Vella
- Tiziano De Matteis
- Zebin Ren
- Luigi Fusco
- Matteo Turisini
- Daniele Cesarini
- Kurt Lust
- Animesh Trivedi
- Duncan Roweth
- Filippo Spiga
- Salvatore Di Girolamo
- Torsten Hoefler
reason: "Comprehensive empirical study of GPU-to-GPU communication across six major supercomputers, revealing bottlenecks and bandwidth characteristics relevant to all distributed AI/HPC workloads."
- title: "Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed\
\ AI"
authors:
- Mikhail Khalilov
- Salvatore Di Girolamo
- Marcin Chrapek
- Rami Nudelman
- Gil Bloch
- Torsten Hoefler
reason: "Achieves bandwidth-optimal collective communication by offloading broadcast and allgather to SmartNICs, directly benefiting large-scale distributed deep learning."
- title: "A Workflow Roofline Model for End-to-End Workflow Performance Analysis"
authors:
- Nan Ding 0006
- Brian Austin
- Yang Liu 0179
- Neil Mehta
- Steven Farrell
- Johannes P. Blaschke
- Leonid Oliker
- Hai Ah Nam
- Nicholas J. Wright
- Samuel Williams 0001
reason: "Extends the Roofline model to full end-to-end HPC workflows, enabling systematic performance diagnosis across compute, I/O, and data movement stages."
- title: "GVARP: Detecting Performance Variance on Large-Scale Heterogeneous Systems"
authors:
- Xin You 0001
- Zhibo Xuan
- Hailong Yang 0002
- Zhongzhi Luan
- Yi Liu 0013
- Depei Qian 0002
reason: "Identifies and diagnoses GPU performance variance at scale on heterogeneous supercomputers, an increasingly critical issue for reproducibility and efficiency."
- title: "A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated\
\ at Exascale"
authors:
- Wesley Brewer
- Matthias Maiterth
- Vineet Kumar
- Rafal P. Wojda
- Sedrick Bouknight
- Jesse Hines
- Woong Shin
- Scott Greenwood
- David Grant
- Wesley Williams
- Feiyi Wang
reason: "First deployment of a digital twin for a liquid-cooled exascale system (Frontier), enabling real-time thermal and power management with validated empirical results."
- title: "Doubling Graph Traversal Efficiency to 198 TeraTEPS on the Supercomputer\
\ Fugaku"
authors:
- Junya Arai
- Masahiro Nakao
- Yuto Inoue
- Kanto Teranishi
- Koji Ueno
- Keiichiro Yamamura
- Mitsuhisa Sato
- Katsuki Fujisawa
reason: "Sets a new world record for graph traversal at 198 TTEPS on Fugaku through novel communication and load-balancing techniques, a landmark Graph500 result."
- title: "MegaMmap: Blurring the Boundary Between Memory and Storage for Data-Intensive\
\ Workloads"
authors:
- Luke Logan
- Anthony Kougkas
- Xian-He Sun
reason: "Novel storage abstraction that transparently tiered memory and storage hierarchies, delivering near-DRAM performance for data-intensive HPC and AI workloads."