--- title: SC 2024 Digest venue: SC year: 2024 date: '2024-01-01' tags: [] paper_count: 15 draft: false --- 15 papers selected. --- ### Pushing the Limit of Quantum Mechanical Simulation to the Raman Spectra of a Biological System with 100 Million Atoms *Honghui Shang, Ying Liu 0055, Zhikun Wu, Zhenchuan Chen *et al.** **TL;DR** — Gordon Bell-class result scaling ab initio Raman spectroscopy to 100 million atoms, pushing quantum-chemical simulation well beyond prior limits. --- ### Breaking the Molecular Dynamics Timescale Barrier Using a Wafer-Scale System *Kylee Santos, Stan G. Moore, Tomas Oppelstrup, Amirali Sharifian *et al.** **TL;DR** — Demonstrates how a Cerebras wafer-scale engine shatters the classical MD timescale barrier, enabling microsecond-regime atomistic simulation at unprecedented speed. --- ### Scaling Molecular Dynamics with ab initio Accuracy to 149 Nanoseconds per Day *Jianxiong Li, Boyang Li, Zhuoqiang Guo, Mingzhen Li 0001 *et al.** **TL;DR** — Achieves 149 ns/day for large-scale deep-potential MD, combining neural-network potentials and HPC engineering to approach DFT accuracy at AIMD-like scale. --- ### Breaking the Million-Electron and 1 EFLOP/s Barriers: Biomolecular-Scale Ab Initio Molecular Dynamics Using MP2 Potentials *Ryan Stocks, Jorge L. Galvez Vallejo, Fiona C. Y. Yu, Calum Snowdon *et al.** **TL;DR** — First demonstration of MP2-level AIMD at the million-electron and exaFLOP/s scale, a landmark in quantum chemistry on supercomputers. --- ### Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning *Wei An, Xiao Bi, Guanting Chen 0002, Shanhuang Chen *et al.** **TL;DR** — Full system co-design report from DeepSeek's AI-HPC cluster showing 40% cost reduction vs. NVIDIA DGX through network/software optimizations, with production evidence at scale. --- ### MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization *Gautham Dharuman, Kyle Hippe, Alexander Brace, Sam Foreman *et al.** **TL;DR** — First exaFLOP/s AI science workflow, integrating multimodal protein design with DPO alignment at supercomputing scale across Frontier and Aurora. --- ### ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability *Xiao Wang 0004, Siyan Liu, Aristeidis Tsaris, Jong-Youl Choi *et al.** **TL;DR** — Introduces a large foundation model for Earth system prediction trained on Frontier, demonstrating how exascale AI infrastructure enables climate-scale spatiotemporal modeling. --- ### Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers *Siddharth Singh, Prajwal Singhania, Aditya K. Ranjan, John Kirchenbauer *et al.** **TL;DR** — Presents an open-source framework for LLM training at thousands-of-GPU scale, systematically analyzing throughput, memory, and communication trade-offs on leadership supercomputers. --- ### Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects *Daniele De Sensi, Lorenzo Pichetti, Flavio Vella, Tiziano De Matteis *et al.** **TL;DR** — Comprehensive empirical study of GPU-to-GPU communication across six major supercomputers, revealing bottlenecks and bandwidth characteristics relevant to all distributed AI/HPC workloads. --- ### Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI *Mikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek, Rami Nudelman *et al.** **TL;DR** — Achieves bandwidth-optimal collective communication by offloading broadcast and allgather to SmartNICs, directly benefiting large-scale distributed deep learning. --- ### A Workflow Roofline Model for End-to-End Workflow Performance Analysis *Nan Ding 0006, Brian Austin, Yang Liu 0179, Neil Mehta *et al.** **TL;DR** — Extends the Roofline model to full end-to-end HPC workflows, enabling systematic performance diagnosis across compute, I/O, and data movement stages. --- ### GVARP: Detecting Performance Variance on Large-Scale Heterogeneous Systems *Xin You 0001, Zhibo Xuan, Hailong Yang 0002, Zhongzhi Luan *et al.** **TL;DR** — Identifies and diagnoses GPU performance variance at scale on heterogeneous supercomputers, an increasingly critical issue for reproducibility and efficiency. --- ### A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated at Exascale *Wesley Brewer, Matthias Maiterth, Vineet Kumar, Rafal P. Wojda *et al.** **TL;DR** — First deployment of a digital twin for a liquid-cooled exascale system (Frontier), enabling real-time thermal and power management with validated empirical results. --- ### Doubling Graph Traversal Efficiency to 198 TeraTEPS on the Supercomputer Fugaku *Junya Arai, Masahiro Nakao, Yuto Inoue, Kanto Teranishi *et al.** **TL;DR** — Sets a new world record for graph traversal at 198 TTEPS on Fugaku through novel communication and load-balancing techniques, a landmark Graph500 result. --- ### MegaMmap: Blurring the Boundary Between Memory and Storage for Data-Intensive Workloads *Luke Logan, Anthony Kougkas, Xian-He Sun* **TL;DR** — Novel storage abstraction that transparently tiered memory and storage hierarchies, delivering near-DRAM performance for data-intensive HPC and AI workloads.