venue: TPDS year: 2024 selected: - title: "Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning" authors: - Burak Aksar - Efe Sencan - Benjamin Schwaller - Omar Aaziz - Vitus J. Leung - Jim M. Brandt - Brian Kulis - Manuel Egele - Ayse K. Coskun tldr: "An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort." why_notable: "It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads." - title: "AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost" authors: - Jinfan Chen - Shigang Li 0002 - Ran Guo - Jinhui Yuan - Torsten Hoefler tldr: "AutoDDL automatically searches for the distributed DNN training strategy that minimizes communication bandwidth cost while meeting performance targets." why_notable: "Combining Hoefler's communication-model expertise with automatic strategy search, this paper is essential reading for practitioners scaling DNN training across large clusters." - title: "PeakFS: An Ultra-High Performance Parallel File System via Computing-Network-Storage Co-Optimization for HPC Applications" authors: - Yixiao Chen - Haomai Yang - Kai Lu 0002 - Wenlve Huang - Jibin Wang - Jiguang Wan 0001 - Jian Zhou 0004 - Fei Wu 0005 - Changsheng Xie 0001 tldr: "PeakFS co-optimizes compute, network, and storage layers of a parallel file system to deliver ultra-high I/O throughput for HPC workloads." why_notable: "Its holistic co-design perspective sets a new performance baseline for HPC storage and provides actionable insights for next-generation parallel file system architects." - title: "Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems" authors: - Chen Wang 0004 - Kathryn M. Mohror - Marc Snir tldr: "This paper formalizes consistency models used by parallel file systems and provides the first systematic empirical comparison of their performance trade-offs." why_notable: "Rigorous formal treatment from Snir and Mohror clarifies long-standing ambiguities in HPC storage semantics, making it an important reference for storage system designers." - title: "Malleability in Modern HPC Systems: Current Experiences, Challenges, and Future Opportunities" authors: - Ahmad Tarraf - Martin Schreiber 0001 - Alberto Cascajo - Jean-Baptiste Besnard - Marc-André Vef - Dominik Huber - Sonja Happ - André Brinkmann - David E. Singh - Hans-Christian Hoppe - Alberto Miranda - Antonio J. Peña - Rui Machado - Marta Garcia-Gasulla - Martin Schulz 0001 - Paul M. Carpenter - Simon Pickartz - Tiberiu Rotaru - Sergio Iserte - Víctor López 0003 - Jorge Ejarque - Heena Sirwani - Jesús Carretero 0001 - Felix Wolf 0001 tldr: "A comprehensive survey of dynamic resource malleability in HPC, covering runtime systems, job schedulers, and application-level support with lessons from production systems." why_notable: "As energy-aware and burst-resilient HPC scheduling becomes critical, this broad community-driven survey is the definitive starting point for research on malleable HPC runtimes." - title: "Pyxis: Scheduling Mixed Tasks in Disaggregated Datacenters" authors: - Sheng Qi - Chao Jin - Mosharaf Chowdhury - Zhenming Liu - Xuanzhe Liu - Xin Jin 0008 tldr: "Pyxis is a scheduler for disaggregated datacenters that jointly manages latency-sensitive and batch tasks by exploiting flexible resource pooling across the disaggregated fabric." why_notable: "It tackles one of the central open problems in cloud scheduling—multi-tenancy under disaggregation—with rigorous analysis and demonstrated gains on real workloads." - title: "Swift: Expedited Failure Recovery for Large-Scale DNN Training" authors: - Yuchen Zhong - Guangming Sheng - Juncheng Liu - Jinhui Yuan - Chuan Wu 0001 tldr: "Swift dramatically reduces checkpoint and recovery overhead for large-scale DNN training by combining lightweight in-memory snapshots with selective recomputation." why_notable: "As training runs on hundreds of GPUs grow longer and failures become inevitable, Swift's fault-tolerance approach directly addresses a practical bottleneck in modern deep-learning infrastructure." - title: "FastLoad: Speeding Up Data Loading of Both Sparse Matrix and Vector for SpMV on GPUs" authors: - Jinyu Hu - Huizhang Luo - Hong Jiang 0001 - Guoqing Xiao 0001 - Kenli Li 0001 tldr: "FastLoad optimizes the memory-access pattern for loading both the sparse matrix and the dense vector in SpMV on GPUs, yielding significant throughput improvements." why_notable: "SpMV is a foundational kernel for scientific computing and graph analytics; this paper's memory-access analysis and optimizations benefit a wide class of GPU applications." - title: "KLNK: Expanding Page Boundaries in a Distributed Shared Memory System" authors: - Yiwei Ci - Michael R. Lyu - Zhan Zhang 0002 - De-Cheng Zuo - Xiao-Zong Yang tldr: "KLNK extends distributed shared memory page granularity to reduce false sharing and improve throughput for irregular access patterns." why_notable: "It addresses a classic but unsolved bottleneck in DSM systems with a practical, page-table-level mechanism applicable to emerging disaggregated memory architectures." - title: "Enabling Efficient Erasure Coding in Disaggregated Memory Systems" authors: - Qiliang Li - Liangliang Xu - Yongkun Li 0001 - Min Lyu - Wei Wang 0502 - Pengfei Zuo - Yinlong Xu 0001 tldr: "This paper designs an erasure-coding scheme tailored to disaggregated memory, exploiting its unique bandwidth topology to achieve fault tolerance with low overhead." why_notable: "Fault tolerance in disaggregated memory is an open problem of growing importance; this work provides concrete mechanisms and strong performance results." - title: "Simple, Fast and Widely Applicable Concurrent Memory Reclamation via Neutralization" authors: - Ajay Singh 0002 - Trevor Alexander Brown - Ali José Mashtizadeh tldr: "Neutralization is a new mechanism for safe memory reclamation in lock-free data structures that is simpler, faster, and more portable than prior approaches." why_notable: "Safe memory reclamation is a pervasive challenge in concurrent programming; this algorithm's breadth of applicability and performance improvements make it highly reusable." - title: "DeepTM: Efficient Tensor Management in Heterogeneous Memory for DNN Training" authors: - Haoran Zhou - Wei Rang - Hongyang Chen 0001 - Xiaobo Zhou 0002 - Dazhao Cheng tldr: "DeepTM dynamically manages tensor placement across DRAM and NVM during DNN training to reduce memory pressure and improve throughput." why_notable: "As model sizes outpace GPU memory, heterogeneous memory management becomes critical; DeepTM provides a practical, training-aware solution with measurable benefits."