All checks were successful
Build and deploy static pages / build-and-push (push) Successful in 19s
133 lines
5.2 KiB
Markdown
133 lines
5.2 KiB
Markdown
---
|
|
title: SC 2024 Digest
|
|
venue: SC
|
|
year: 2024
|
|
date: '2024-01-01'
|
|
tags: []
|
|
paper_count: 15
|
|
draft: false
|
|
---
|
|
|
|
15 papers selected.
|
|
|
|
---
|
|
|
|
### Pushing the Limit of Quantum Mechanical Simulation to the Raman Spectra of a Biological System with 100 Million Atoms
|
|
|
|
*Honghui Shang, Ying Liu 0055, Zhikun Wu, Zhenchuan Chen *et al.**
|
|
|
|
**TL;DR** — Gordon Bell-class result scaling ab initio Raman spectroscopy to 100 million atoms, pushing quantum-chemical simulation well beyond prior limits.
|
|
|
|
---
|
|
|
|
### Breaking the Molecular Dynamics Timescale Barrier Using a Wafer-Scale System
|
|
|
|
*Kylee Santos, Stan G. Moore, Tomas Oppelstrup, Amirali Sharifian *et al.**
|
|
|
|
**TL;DR** — Demonstrates how a Cerebras wafer-scale engine shatters the classical MD timescale barrier, enabling microsecond-regime atomistic simulation at unprecedented speed.
|
|
|
|
---
|
|
|
|
### Scaling Molecular Dynamics with ab initio Accuracy to 149 Nanoseconds per Day
|
|
|
|
*Jianxiong Li, Boyang Li, Zhuoqiang Guo, Mingzhen Li 0001 *et al.**
|
|
|
|
**TL;DR** — Achieves 149 ns/day for large-scale deep-potential MD, combining neural-network potentials and HPC engineering to approach DFT accuracy at AIMD-like scale.
|
|
|
|
---
|
|
|
|
### Breaking the Million-Electron and 1 EFLOP/s Barriers: Biomolecular-Scale Ab Initio Molecular Dynamics Using MP2 Potentials
|
|
|
|
*Ryan Stocks, Jorge L. Galvez Vallejo, Fiona C. Y. Yu, Calum Snowdon *et al.**
|
|
|
|
**TL;DR** — First demonstration of MP2-level AIMD at the million-electron and exaFLOP/s scale, a landmark in quantum chemistry on supercomputers.
|
|
|
|
---
|
|
|
|
### Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
|
|
|
|
*Wei An, Xiao Bi, Guanting Chen 0002, Shanhuang Chen *et al.**
|
|
|
|
**TL;DR** — Full system co-design report from DeepSeek's AI-HPC cluster showing 40% cost reduction vs. NVIDIA DGX through network/software optimizations, with production evidence at scale.
|
|
|
|
---
|
|
|
|
### MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization
|
|
|
|
*Gautham Dharuman, Kyle Hippe, Alexander Brace, Sam Foreman *et al.**
|
|
|
|
**TL;DR** — First exaFLOP/s AI science workflow, integrating multimodal protein design with DPO alignment at supercomputing scale across Frontier and Aurora.
|
|
|
|
---
|
|
|
|
### ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability
|
|
|
|
*Xiao Wang 0004, Siyan Liu, Aristeidis Tsaris, Jong-Youl Choi *et al.**
|
|
|
|
**TL;DR** — Introduces a large foundation model for Earth system prediction trained on Frontier, demonstrating how exascale AI infrastructure enables climate-scale spatiotemporal modeling.
|
|
|
|
---
|
|
|
|
### Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
|
|
|
|
*Siddharth Singh, Prajwal Singhania, Aditya K. Ranjan, John Kirchenbauer *et al.**
|
|
|
|
**TL;DR** — Presents an open-source framework for LLM training at thousands-of-GPU scale, systematically analyzing throughput, memory, and communication trade-offs on leadership supercomputers.
|
|
|
|
---
|
|
|
|
### Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
|
|
|
|
*Daniele De Sensi, Lorenzo Pichetti, Flavio Vella, Tiziano De Matteis *et al.**
|
|
|
|
**TL;DR** — Comprehensive empirical study of GPU-to-GPU communication across six major supercomputers, revealing bottlenecks and bandwidth characteristics relevant to all distributed AI/HPC workloads.
|
|
|
|
---
|
|
|
|
### Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
|
|
|
|
*Mikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek, Rami Nudelman *et al.**
|
|
|
|
**TL;DR** — Achieves bandwidth-optimal collective communication by offloading broadcast and allgather to SmartNICs, directly benefiting large-scale distributed deep learning.
|
|
|
|
---
|
|
|
|
### A Workflow Roofline Model for End-to-End Workflow Performance Analysis
|
|
|
|
*Nan Ding 0006, Brian Austin, Yang Liu 0179, Neil Mehta *et al.**
|
|
|
|
**TL;DR** — Extends the Roofline model to full end-to-end HPC workflows, enabling systematic performance diagnosis across compute, I/O, and data movement stages.
|
|
|
|
---
|
|
|
|
### GVARP: Detecting Performance Variance on Large-Scale Heterogeneous Systems
|
|
|
|
*Xin You 0001, Zhibo Xuan, Hailong Yang 0002, Zhongzhi Luan *et al.**
|
|
|
|
**TL;DR** — Identifies and diagnoses GPU performance variance at scale on heterogeneous supercomputers, an increasingly critical issue for reproducibility and efficiency.
|
|
|
|
---
|
|
|
|
### A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated at Exascale
|
|
|
|
*Wesley Brewer, Matthias Maiterth, Vineet Kumar, Rafal P. Wojda *et al.**
|
|
|
|
**TL;DR** — First deployment of a digital twin for a liquid-cooled exascale system (Frontier), enabling real-time thermal and power management with validated empirical results.
|
|
|
|
---
|
|
|
|
### Doubling Graph Traversal Efficiency to 198 TeraTEPS on the Supercomputer Fugaku
|
|
|
|
*Junya Arai, Masahiro Nakao, Yuto Inoue, Kanto Teranishi *et al.**
|
|
|
|
**TL;DR** — Sets a new world record for graph traversal at 198 TTEPS on Fugaku through novel communication and load-balancing techniques, a landmark Graph500 result.
|
|
|
|
---
|
|
|
|
### MegaMmap: Blurring the Boundary Between Memory and Storage for Data-Intensive Workloads
|
|
|
|
*Luke Logan, Anthony Kougkas, Xian-He Sun*
|
|
|
|
**TL;DR** — Novel storage abstraction that transparently tiered memory and storage hierarchies, delivering near-DRAM performance for data-intensive HPC and AI workloads.
|
|
|