Files
publish-assistant/site/content/digests/SC-2024/index.md
2026-04-26 12:57:40 +00:00

5.2 KiB

title, venue, year, date, tags, paper_count, draft
title venue year date tags paper_count draft
SC 2024 Digest SC 2024 2024-01-01
15 false

15 papers selected.


Pushing the Limit of Quantum Mechanical Simulation to the Raman Spectra of a Biological System with 100 Million Atoms

Honghui Shang, Ying Liu 0055, Zhikun Wu, Zhenchuan Chen et al.

TL;DR — Gordon Bell-class result scaling ab initio Raman spectroscopy to 100 million atoms, pushing quantum-chemical simulation well beyond prior limits.


Breaking the Molecular Dynamics Timescale Barrier Using a Wafer-Scale System

Kylee Santos, Stan G. Moore, Tomas Oppelstrup, Amirali Sharifian et al.

TL;DR — Demonstrates how a Cerebras wafer-scale engine shatters the classical MD timescale barrier, enabling microsecond-regime atomistic simulation at unprecedented speed.


Scaling Molecular Dynamics with ab initio Accuracy to 149 Nanoseconds per Day

Jianxiong Li, Boyang Li, Zhuoqiang Guo, Mingzhen Li 0001 et al.

TL;DR — Achieves 149 ns/day for large-scale deep-potential MD, combining neural-network potentials and HPC engineering to approach DFT accuracy at AIMD-like scale.


Breaking the Million-Electron and 1 EFLOP/s Barriers: Biomolecular-Scale Ab Initio Molecular Dynamics Using MP2 Potentials

Ryan Stocks, Jorge L. Galvez Vallejo, Fiona C. Y. Yu, Calum Snowdon et al.

TL;DR — First demonstration of MP2-level AIMD at the million-electron and exaFLOP/s scale, a landmark in quantum chemistry on supercomputers.


Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning

Wei An, Xiao Bi, Guanting Chen 0002, Shanhuang Chen et al.

TL;DR — Full system co-design report from DeepSeek's AI-HPC cluster showing 40% cost reduction vs. NVIDIA DGX through network/software optimizations, with production evidence at scale.


MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization

Gautham Dharuman, Kyle Hippe, Alexander Brace, Sam Foreman et al.

TL;DR — First exaFLOP/s AI science workflow, integrating multimodal protein design with DPO alignment at supercomputing scale across Frontier and Aurora.


ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability

Xiao Wang 0004, Siyan Liu, Aristeidis Tsaris, Jong-Youl Choi et al.

TL;DR — Introduces a large foundation model for Earth system prediction trained on Frontier, demonstrating how exascale AI infrastructure enables climate-scale spatiotemporal modeling.


Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers

Siddharth Singh, Prajwal Singhania, Aditya K. Ranjan, John Kirchenbauer et al.

TL;DR — Presents an open-source framework for LLM training at thousands-of-GPU scale, systematically analyzing throughput, memory, and communication trade-offs on leadership supercomputers.


Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects

Daniele De Sensi, Lorenzo Pichetti, Flavio Vella, Tiziano De Matteis et al.

TL;DR — Comprehensive empirical study of GPU-to-GPU communication across six major supercomputers, revealing bottlenecks and bandwidth characteristics relevant to all distributed AI/HPC workloads.


Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI

Mikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek, Rami Nudelman et al.

TL;DR — Achieves bandwidth-optimal collective communication by offloading broadcast and allgather to SmartNICs, directly benefiting large-scale distributed deep learning.


A Workflow Roofline Model for End-to-End Workflow Performance Analysis

Nan Ding 0006, Brian Austin, Yang Liu 0179, Neil Mehta et al.

TL;DR — Extends the Roofline model to full end-to-end HPC workflows, enabling systematic performance diagnosis across compute, I/O, and data movement stages.


GVARP: Detecting Performance Variance on Large-Scale Heterogeneous Systems

Xin You 0001, Zhibo Xuan, Hailong Yang 0002, Zhongzhi Luan et al.

TL;DR — Identifies and diagnoses GPU performance variance at scale on heterogeneous supercomputers, an increasingly critical issue for reproducibility and efficiency.


A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated at Exascale

Wesley Brewer, Matthias Maiterth, Vineet Kumar, Rafal P. Wojda et al.

TL;DR — First deployment of a digital twin for a liquid-cooled exascale system (Frontier), enabling real-time thermal and power management with validated empirical results.


Doubling Graph Traversal Efficiency to 198 TeraTEPS on the Supercomputer Fugaku

Junya Arai, Masahiro Nakao, Yuto Inoue, Kanto Teranishi et al.

TL;DR — Sets a new world record for graph traversal at 198 TTEPS on Fugaku through novel communication and load-balancing techniques, a landmark Graph500 result.


MegaMmap: Blurring the Boundary Between Memory and Storage for Data-Intensive Workloads

Luke Logan, Anthony Kougkas, Xian-He Sun

TL;DR — Novel storage abstraction that transparently tiered memory and storage hierarchies, delivering near-DRAM performance for data-intensive HPC and AI workloads.