Files
2026-08-18 13:39:21 +00:00

29 lines
27 KiB
HTML
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!doctype html><html lang=en dir=auto data-theme=auto><head><meta charset=utf-8><meta http-equiv=X-UA-Compatible content="IE=edge"><meta name=viewport content="width=device-width,initial-scale=1,shrink-to-fit=no"><meta name=robots content="index, follow"><title>TPDS 2025 Digest | Publish Assistant</title><meta name=keywords content><meta name=description content="12 papers selected.
HARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management
Kyrian Adimora, Hongyang Sun 0001
TL;DR — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.
Why notable — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.
MIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems
Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie et al."><meta name=author content="Publish Assistant"><link rel=canonical href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2025/><link crossorigin=anonymous href=/vincent/publish-assistant/assets/css/stylesheet.d72f07832e13c592b3edba91680bfe70f01daac396179bcace0ac36e8e0494c6.css integrity="sha256-1y8Hgy4TxZKz7bqRaAv+cPAdqsOWF5vKzgrDbo4ElMY=" rel="preload stylesheet" as=style><link rel=icon href=https://pub.sqrt.fr/vincent/publish-assistant/favicon.ico><link rel=icon type=image/png sizes=16x16 href=https://pub.sqrt.fr/vincent/publish-assistant/favicon-16x16.png><link rel=icon type=image/png sizes=32x32 href=https://pub.sqrt.fr/vincent/publish-assistant/favicon-32x32.png><link rel=apple-touch-icon href=https://pub.sqrt.fr/vincent/publish-assistant/apple-touch-icon.png><link rel=mask-icon href=https://pub.sqrt.fr/vincent/publish-assistant/safari-pinned-tab.svg><meta name=theme-color content="#2e2e33"><meta name=msapplication-TileColor content="#2e2e33"><link rel=alternate hreflang=en href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2025/><noscript><style>#theme-toggle,.top-link{display:none}</style><style>@media(prefers-color-scheme:dark){:root{--theme:rgb(29, 30, 32);--entry:rgb(46, 46, 51);--primary:rgb(218, 218, 219);--secondary:rgb(155, 156, 157);--tertiary:rgb(65, 66, 68);--content:rgb(196, 196, 197);--code-block-bg:rgb(46, 46, 51);--code-bg:rgb(55, 56, 62);--border:rgb(51, 51, 51);color-scheme:dark}.list{background:var(--theme)}.toc{background:var(--entry)}}</style></noscript><script>localStorage.getItem("pref-theme")==="dark"?document.querySelector("html").dataset.theme="dark":localStorage.getItem("pref-theme")==="light"?document.querySelector("html").dataset.theme="light":window.matchMedia("(prefers-color-scheme: dark)").matches?document.querySelector("html").dataset.theme="dark":document.querySelector("html").dataset.theme="light"</script><meta property="og:url" content="https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2025/"><meta property="og:site_name" content="Publish Assistant"><meta property="og:title" content="TPDS 2025 Digest"><meta property="og:description" content="12 papers selected.
HARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management Kyrian Adimora, Hongyang Sun 0001
TL;DR — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.
Why notable — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.
MIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie et al."><meta property="og:locale" content="en_us"><meta property="og:type" content="article"><meta property="article:section" content="cloud-edge"><meta property="article:published_time" content="2025-01-01T00:00:00+00:00"><meta property="article:modified_time" content="2025-01-01T00:00:00+00:00"><meta name=twitter:card content="summary"><meta name=twitter:title content="TPDS 2025 Digest"><meta name=twitter:description content="12 papers selected.
HARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management Kyrian Adimora, Hongyang Sun 0001
TL;DR — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.
Why notable — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.
MIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie et al."><script type=application/ld+json>{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Edge and Cloud Systems","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/"},{"@type":"ListItem","position":2,"name":"Digests","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/"},{"@type":"ListItem","position":3,"name":"TPDS 2025 Digest","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2025/"}]}</script><script type=application/ld+json>{"@context":"https://schema.org","@type":"BlogPosting","headline":"TPDS 2025 Digest","name":"TPDS 2025 Digest","description":"12 papers selected.\nHARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management Kyrian Adimora, Hongyang Sun 0001\nTL;DR — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.\nWhy notable — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.\nMIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie et al.\n","keywords":[],"articleBody":"12 papers selected.\nHARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management Kyrian Adimora, Hongyang Sun 0001\nTL;DR — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.\nWhy notable — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.\nMIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie et al.\nTL;DR — MIST reduces MPI job startup and termination latency to near-instant on the Tianhe supercomputer by redesigning the process-management and communication-bootstrap path.\nWhy notable — Startup overhead is a significant fraction of short-job turnaround time at scale; MISTs results on a top-ranked system provide a concrete reference for HPC runtime developers.\nScheduling With Lightweight Predictions in Power-Constrained HPC Platforms Danilo Carastan-Santos, Georges Da Costa, Igor Fontana De Nardin, Millian Poquet et al.\nTL;DR — This paper develops a scheduling framework that uses lightweight runtime predictions to respect power caps on HPC systems while minimizing job slowdown.\nWhy notable — Power capping is now a first-class constraint on modern supercomputers, and this work from leading European HPC scheduling researchers offers practical, deployable algorithms.\nPipeMesh: Achieving Memory-Efficient Computation-Communication Overlap for Training Large Language Models Fanxin Li, Shixiong Zhao, Yuhao Qing, Jianyu Jiang et al.\nTL;DR — PipeMesh overlaps pipeline-parallel computation and communication for LLM training while carefully managing memory to avoid out-of-memory failures.\nWhy notable — Communication-computation overlap is one of the most impactful levers for LLM training efficiency, and PipeMeshs memory-awareness addresses the key practical constraint.\nEfficientMoE: Optimizing Mixture-of-Experts Model Training With Adaptive Load Balance Yan Zeng, Chengchuang Huang, Yipeng Mei, Lifu Zhang 0004 et al.\nTL;DR — EfficientMoE introduces an adaptive load-balancing strategy for Mixture-of-Experts training that equalizes expert utilization and reduces communication bottlenecks.\nWhy notable — MoE models are central to frontier LLM architectures, and load imbalance is their primary training inefficiency; this work provides both analysis and a practical solution.\nSSpMM: Efficiently Scalable SpMM Kernels Across Multiple Generations of Tensor Cores Zeyu Xue, Mei Wen, Jianchao Yang, Minjin Tang et al.\nTL;DR — SSpMM delivers portable, high-performance sparse-matrix dense-matrix multiplication kernels that scale efficiently across Ampere, Hopper, and future Tensor Core generations.\nWhy notable — SpMM is a bottleneck in GNN training and scientific computing; cross-generation portability without performance loss is a significant contribution for the GPU computing community.\nIceFrog: A Layer-Elastic Scheduling System for Deep Learning Training in GPU Clusters Wei Gao 0064, Zhuoyuan Ouyang, Peng Sun 0006, Tianwei Zhang 0004 et al.\nTL;DR — IceFrog dynamically adjusts the number of pipeline stages (layers) assigned to each GPU during training to adapt to cluster heterogeneity and improve utilization.\nWhy notable — Layer elasticity is a novel dimension of flexibility in distributed DNN training; IceFrogs scheduler provides measurable throughput gains in realistic heterogeneous GPU clusters.\nElastic Relaxation of Concurrent Data Structures Kåre von Geijer, Philippas Tsigas\nTL;DR — This paper introduces a formal framework and concrete algorithms for elastic relaxation of concurrent data structures, allowing tunable trade-offs between consistency and throughput.\nWhy notable — Tsigass group advances concurrent data-structure theory with a unifying formalism that subsumes many ad-hoc relaxed designs and enables provable guarantees.\nApproximation Algorithms for Scheduling With/Without Deadline Constraints Where Rejection Costs are Proportional to Processing Times Olivier Beaumont, Rémi Bouzel, Lionel Eyraud-Dubois, Esragul Korkmaz et al.\nTL;DR — This paper derives new approximation algorithms with tight ratios for online and offline scheduling problems where rejected jobs incur costs proportional to their processing times.\nWhy notable — The theoretical results close open gaps in parallel scheduling complexity and are directly applicable to cloud and HPC batch schedulers that must handle job rejection.\nEdgeHydra: Fault-Tolerant Edge Data Distribution Based on Erasure Coding Qiang He 0001, Guobiao Zhang, Jiawei Wang 0003, Ruikun Luo et al.\nTL;DR — EdgeHydra uses erasure coding tailored to edge-node failure patterns to provide fault-tolerant data distribution with low redundancy overhead at the network edge.\nWhy notable — Fault tolerance at the edge is an increasingly critical requirement, and this systems erasure-coding approach significantly outperforms replication in storage efficiency.\nTwo-Dimensional Balanced Partitioning and Efficient Caching for Distributed Graph Analysis Shuai Lin, Rui Wang 0076, Yongkun Li 0001, Yinlong Xu 0001 et al.\nTL;DR — This paper proposes a 2D balanced graph partitioning scheme combined with a caching policy that jointly minimizes communication and replication costs in distributed graph systems.\nWhy notable — Graph partitioning and caching are co-dependent problems rarely treated together; the combined optimization yields substantial performance improvements with strong theoretical backing.\nTowards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms Zhongyi Lin, Ning Sun, Pallab Bhattacharya, Xizhou Feng et al.\nTL;DR — This work builds a platform-agnostic performance model for distributed ML training that accurately predicts training throughput across diverse multi-GPU configurations without per-system profiling.\nWhy notable — A universal modeling framework from Owenss group removes the need for expensive empirical searches when tuning distributed training configurations, benefiting the entire ML systems community.\n","wordCount":"838","inLanguage":"en","datePublished":"2025-01-01T00:00:00Z","dateModified":"2025-01-01T00:00:00Z","author":{"@type":"Person","name":"Publish Assistant"},"mainEntityOfPage":{"@type":"WebPage","@id":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2025/"},"publisher":{"@type":"Organization","name":"Publish Assistant","logo":{"@type":"ImageObject","url":"https://pub.sqrt.fr/vincent/publish-assistant/favicon.ico"}}}</script></head><body id=top><header class=header><nav class=header-nav><div class=logo><a href=https://pub.sqrt.fr/vincent/publish-assistant/ accesskey=h title="Publish Assistant (Alt + H)">Publish Assistant</a>
<span class=logo-sep>/</span>
<a class=logo-topic href=/vincent/publish-assistant/cloud-edge/ title="Edge and Cloud Systems">Edge and Cloud Systems</a><div class=logo-switches><button id=theme-toggle class=theme-toggle accesskey=t title="(Alt + T)" aria-label="Toggle theme">
<svg class="moon" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1111.21 3 7 7 0 0021 12.79z"/></svg>
<svg class="sun" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg></button></div></div><ul id=menu class=menu><li><a href=/vincent/publish-assistant/cloud-edge/venues/ title=Venues><span>Venues</span></a></li><li><a href=/vincent/publish-assistant/cloud-edge/calendar/ title=Calendar><span>Calendar</span></a></li><li><a href=/vincent/publish-assistant/cloud-edge/digests/ title=Digests><span class=active>Digests</span></a></li></ul></nav></header><main class=main><article class=post-single><header class=post-header><nav class=breadcrumbs role=navigation aria-label=Breadcrumb><a href=/vincent/publish-assistant/cloud-edge/digests/>Digests</a>
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevron-right"><polyline points="9 18 15 12 9 6"/></svg></nav><h1 class="post-title entry-hint-parent">TPDS 2025 Digest</h1><div class=post-meta><span title='2025-01-01 00:00:00 +0000 UTC'>January 1, 2025</span>&nbsp;·&nbsp;<span>Publish Assistant</span></div></header><div class="post-content md-content"><p>12 papers selected.</p><hr><h3 id=harmonic-uncertainty-aware-multi-objective-optimization-for-energy-efficient-hpc-resource-management>HARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management<a hidden class=anchor aria-hidden=true href=#harmonic-uncertainty-aware-multi-objective-optimization-for-energy-efficient-hpc-resource-management>#</a></h3><p><em>Kyrian Adimora, Hongyang Sun 0001</em></p><p><strong>TL;DR</strong> — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.</p><p><strong>Why notable</strong> — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.</p><hr><h3 id=mist-towards-mpi-instant-startup-and-termination-on-tianhe-hpc-systems>MIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems<a hidden class=anchor aria-hidden=true href=#mist-towards-mpi-instant-startup-and-termination-on-tianhe-hpc-systems>#</a></h3><p><em>Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie <em>et al.</em></em></p><p><strong>TL;DR</strong> — MIST reduces MPI job startup and termination latency to near-instant on the Tianhe supercomputer by redesigning the process-management and communication-bootstrap path.</p><p><strong>Why notable</strong> — Startup overhead is a significant fraction of short-job turnaround time at scale; MIST&rsquo;s results on a top-ranked system provide a concrete reference for HPC runtime developers.</p><hr><h3 id=scheduling-with-lightweight-predictions-in-power-constrained-hpc-platforms>Scheduling With Lightweight Predictions in Power-Constrained HPC Platforms<a hidden class=anchor aria-hidden=true href=#scheduling-with-lightweight-predictions-in-power-constrained-hpc-platforms>#</a></h3><p><em>Danilo Carastan-Santos, Georges Da Costa, Igor Fontana De Nardin, Millian Poquet <em>et al.</em></em></p><p><strong>TL;DR</strong> — This paper develops a scheduling framework that uses lightweight runtime predictions to respect power caps on HPC systems while minimizing job slowdown.</p><p><strong>Why notable</strong> — Power capping is now a first-class constraint on modern supercomputers, and this work from leading European HPC scheduling researchers offers practical, deployable algorithms.</p><hr><h3 id=pipemesh-achieving-memory-efficient-computation-communication-overlap-for-training-large-language-models>PipeMesh: Achieving Memory-Efficient Computation-Communication Overlap for Training Large Language Models<a hidden class=anchor aria-hidden=true href=#pipemesh-achieving-memory-efficient-computation-communication-overlap-for-training-large-language-models>#</a></h3><p><em>Fanxin Li, Shixiong Zhao, Yuhao Qing, Jianyu Jiang <em>et al.</em></em></p><p><strong>TL;DR</strong> — PipeMesh overlaps pipeline-parallel computation and communication for LLM training while carefully managing memory to avoid out-of-memory failures.</p><p><strong>Why notable</strong> — Communication-computation overlap is one of the most impactful levers for LLM training efficiency, and PipeMesh&rsquo;s memory-awareness addresses the key practical constraint.</p><hr><h3 id=efficientmoe-optimizing-mixture-of-experts-model-training-with-adaptive-load-balance>EfficientMoE: Optimizing Mixture-of-Experts Model Training With Adaptive Load Balance<a hidden class=anchor aria-hidden=true href=#efficientmoe-optimizing-mixture-of-experts-model-training-with-adaptive-load-balance>#</a></h3><p><em>Yan Zeng, Chengchuang Huang, Yipeng Mei, Lifu Zhang 0004 <em>et al.</em></em></p><p><strong>TL;DR</strong> — EfficientMoE introduces an adaptive load-balancing strategy for Mixture-of-Experts training that equalizes expert utilization and reduces communication bottlenecks.</p><p><strong>Why notable</strong> — MoE models are central to frontier LLM architectures, and load imbalance is their primary training inefficiency; this work provides both analysis and a practical solution.</p><hr><h3 id=sspmm-efficiently-scalable-spmm-kernels-across-multiple-generations-of-tensor-cores>SSpMM: Efficiently Scalable SpMM Kernels Across Multiple Generations of Tensor Cores<a hidden class=anchor aria-hidden=true href=#sspmm-efficiently-scalable-spmm-kernels-across-multiple-generations-of-tensor-cores>#</a></h3><p><em>Zeyu Xue, Mei Wen, Jianchao Yang, Minjin Tang <em>et al.</em></em></p><p><strong>TL;DR</strong> — SSpMM delivers portable, high-performance sparse-matrix dense-matrix multiplication kernels that scale efficiently across Ampere, Hopper, and future Tensor Core generations.</p><p><strong>Why notable</strong> — SpMM is a bottleneck in GNN training and scientific computing; cross-generation portability without performance loss is a significant contribution for the GPU computing community.</p><hr><h3 id=icefrog-a-layer-elastic-scheduling-system-for-deep-learning-training-in-gpu-clusters>IceFrog: A Layer-Elastic Scheduling System for Deep Learning Training in GPU Clusters<a hidden class=anchor aria-hidden=true href=#icefrog-a-layer-elastic-scheduling-system-for-deep-learning-training-in-gpu-clusters>#</a></h3><p><em>Wei Gao 0064, Zhuoyuan Ouyang, Peng Sun 0006, Tianwei Zhang 0004 <em>et al.</em></em></p><p><strong>TL;DR</strong> — IceFrog dynamically adjusts the number of pipeline stages (layers) assigned to each GPU during training to adapt to cluster heterogeneity and improve utilization.</p><p><strong>Why notable</strong> — Layer elasticity is a novel dimension of flexibility in distributed DNN training; IceFrog&rsquo;s scheduler provides measurable throughput gains in realistic heterogeneous GPU clusters.</p><hr><h3 id=elastic-relaxation-of-concurrent-data-structures>Elastic Relaxation of Concurrent Data Structures<a hidden class=anchor aria-hidden=true href=#elastic-relaxation-of-concurrent-data-structures>#</a></h3><p><em>Kåre von Geijer, Philippas Tsigas</em></p><p><strong>TL;DR</strong> — This paper introduces a formal framework and concrete algorithms for elastic relaxation of concurrent data structures, allowing tunable trade-offs between consistency and throughput.</p><p><strong>Why notable</strong> — Tsigas&rsquo;s group advances concurrent data-structure theory with a unifying formalism that subsumes many ad-hoc relaxed designs and enables provable guarantees.</p><hr><h3 id=approximation-algorithms-for-scheduling-withwithout-deadline-constraints-where-rejection-costs-are-proportional-to-processing-times>Approximation Algorithms for Scheduling With/Without Deadline Constraints Where Rejection Costs are Proportional to Processing Times<a hidden class=anchor aria-hidden=true href=#approximation-algorithms-for-scheduling-withwithout-deadline-constraints-where-rejection-costs-are-proportional-to-processing-times>#</a></h3><p><em>Olivier Beaumont, Rémi Bouzel, Lionel Eyraud-Dubois, Esragul Korkmaz <em>et al.</em></em></p><p><strong>TL;DR</strong> — This paper derives new approximation algorithms with tight ratios for online and offline scheduling problems where rejected jobs incur costs proportional to their processing times.</p><p><strong>Why notable</strong> — The theoretical results close open gaps in parallel scheduling complexity and are directly applicable to cloud and HPC batch schedulers that must handle job rejection.</p><hr><h3 id=edgehydra-fault-tolerant-edge-data-distribution-based-on-erasure-coding>EdgeHydra: Fault-Tolerant Edge Data Distribution Based on Erasure Coding<a hidden class=anchor aria-hidden=true href=#edgehydra-fault-tolerant-edge-data-distribution-based-on-erasure-coding>#</a></h3><p><em>Qiang He 0001, Guobiao Zhang, Jiawei Wang 0003, Ruikun Luo <em>et al.</em></em></p><p><strong>TL;DR</strong> — EdgeHydra uses erasure coding tailored to edge-node failure patterns to provide fault-tolerant data distribution with low redundancy overhead at the network edge.</p><p><strong>Why notable</strong> — Fault tolerance at the edge is an increasingly critical requirement, and this system&rsquo;s erasure-coding approach significantly outperforms replication in storage efficiency.</p><hr><h3 id=two-dimensional-balanced-partitioning-and-efficient-caching-for-distributed-graph-analysis>Two-Dimensional Balanced Partitioning and Efficient Caching for Distributed Graph Analysis<a hidden class=anchor aria-hidden=true href=#two-dimensional-balanced-partitioning-and-efficient-caching-for-distributed-graph-analysis>#</a></h3><p><em>Shuai Lin, Rui Wang 0076, Yongkun Li 0001, Yinlong Xu 0001 <em>et al.</em></em></p><p><strong>TL;DR</strong> — This paper proposes a 2D balanced graph partitioning scheme combined with a caching policy that jointly minimizes communication and replication costs in distributed graph systems.</p><p><strong>Why notable</strong> — Graph partitioning and caching are co-dependent problems rarely treated together; the combined optimization yields substantial performance improvements with strong theoretical backing.</p><hr><h3 id=towards-universal-performance-modeling-for-machine-learning-training-on-multi-gpu-platforms>Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms<a hidden class=anchor aria-hidden=true href=#towards-universal-performance-modeling-for-machine-learning-training-on-multi-gpu-platforms>#</a></h3><p><em>Zhongyi Lin, Ning Sun, Pallab Bhattacharya, Xizhou Feng <em>et al.</em></em></p><p><strong>TL;DR</strong> — This work builds a platform-agnostic performance model for distributed ML training that accurately predicts training throughput across diverse multi-GPU configurations without per-system profiling.</p><p><strong>Why notable</strong> — A universal modeling framework from Owens&rsquo;s group removes the need for expensive empirical searches when tuning distributed training configurations, benefiting the entire ML systems community.</p></div><footer class=post-footer><ul class=post-tags></ul><nav class=paginav><a class=prev href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tocs-2025/><span class=title>« Prev</span>
<span>TOCS 2025 Digest</span>
</a><a class=next href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/middleware-2024/><span class=title>Next »</span>
<span>Middleware 2024 Digest</span></a></nav></footer></article></main><footer class=footer><span>&copy; 2026 <a href=https://pub.sqrt.fr/vincent/publish-assistant/>Publish Assistant</a></span> ·
<span>Powered by
<a href="https://gohugo.io/?utm_source=papermod" rel=noopener target=_blank>Hugo</a> &
<a href=https://github.com/adityatelange/hugo-PaperMod/ rel=noopener target=_blank>PaperMod</a></span></footer><a href=#top id=top-link class="top-link hidden" aria-label="go to top" title="Go to Top (Alt + G)" accesskey=g><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevrons-up"><polyline points="17 11 12 6 7 11"/><polyline points="17 18 12 13 7 18"/></svg>
</a><script>let menu=document.getElementById("menu");if(menu){const e=localStorage.getItem("menu-scroll-position");e&&(menu.scrollLeft=parseInt(e,10)),menu.onscroll=function(){localStorage.setItem("menu-scroll-position",menu.scrollLeft)}}document.querySelectorAll('a[href^="#"]').forEach(e=>{e.addEventListener("click",function(e){e.preventDefault();var t=this.getAttribute("href").substr(1);window.matchMedia("(prefers-reduced-motion: reduce)").matches?document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView():document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView({behavior:"smooth"}),t==="top"?history.replaceState(null,null," "):history.pushState(null,null,`#${t}`)})})</script><script>var toplink=document.getElementById("top-link");window.onscroll=function(){const e=window.innerHeight;document.body.scrollTop>e||document.documentElement.scrollTop>e?toplink.classList.remove("hidden"):toplink.classList.add("hidden")}</script><script>document.getElementById("theme-toggle").addEventListener("click",()=>{const e=document.querySelector("html");e.dataset.theme==="dark"?(e.dataset.theme="light",localStorage.setItem("pref-theme","light")):(e.dataset.theme="dark",localStorage.setItem("pref-theme","dark"))})</script></body></html>