29 lines
27 KiB
HTML
29 lines
27 KiB
HTML
<!doctype html><html lang=en dir=auto data-theme=auto><head><meta charset=utf-8><meta http-equiv=X-UA-Compatible content="IE=edge"><meta name=viewport content="width=device-width,initial-scale=1,shrink-to-fit=no"><meta name=robots content="index, follow"><title>TPDS 2025 Digest | Publish Assistant</title><meta name=keywords content><meta name=description content="12 papers selected.
|
||
|
||
HARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management
|
||
Kyrian Adimora, Hongyang Sun 0001
|
||
TL;DR — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.
|
||
Why notable — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.
|
||
|
||
MIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems
|
||
Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie et al."><meta name=author content="Publish Assistant"><link rel=canonical href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2025/><link crossorigin=anonymous href=/vincent/publish-assistant/assets/css/stylesheet.d72f07832e13c592b3edba91680bfe70f01daac396179bcace0ac36e8e0494c6.css integrity="sha256-1y8Hgy4TxZKz7bqRaAv+cPAdqsOWF5vKzgrDbo4ElMY=" rel="preload stylesheet" as=style><link rel=icon href=https://pub.sqrt.fr/vincent/publish-assistant/favicon.ico><link rel=icon type=image/png sizes=16x16 href=https://pub.sqrt.fr/vincent/publish-assistant/favicon-16x16.png><link rel=icon type=image/png sizes=32x32 href=https://pub.sqrt.fr/vincent/publish-assistant/favicon-32x32.png><link rel=apple-touch-icon href=https://pub.sqrt.fr/vincent/publish-assistant/apple-touch-icon.png><link rel=mask-icon href=https://pub.sqrt.fr/vincent/publish-assistant/safari-pinned-tab.svg><meta name=theme-color content="#2e2e33"><meta name=msapplication-TileColor content="#2e2e33"><link rel=alternate hreflang=en href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2025/><noscript><style>#theme-toggle,.top-link{display:none}</style><style>@media(prefers-color-scheme:dark){:root{--theme:rgb(29, 30, 32);--entry:rgb(46, 46, 51);--primary:rgb(218, 218, 219);--secondary:rgb(155, 156, 157);--tertiary:rgb(65, 66, 68);--content:rgb(196, 196, 197);--code-block-bg:rgb(46, 46, 51);--code-bg:rgb(55, 56, 62);--border:rgb(51, 51, 51);color-scheme:dark}.list{background:var(--theme)}.toc{background:var(--entry)}}</style></noscript><script>localStorage.getItem("pref-theme")==="dark"?document.querySelector("html").dataset.theme="dark":localStorage.getItem("pref-theme")==="light"?document.querySelector("html").dataset.theme="light":window.matchMedia("(prefers-color-scheme: dark)").matches?document.querySelector("html").dataset.theme="dark":document.querySelector("html").dataset.theme="light"</script><meta property="og:url" content="https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2025/"><meta property="og:site_name" content="Publish Assistant"><meta property="og:title" content="TPDS 2025 Digest"><meta property="og:description" content="12 papers selected.
|
||
HARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management Kyrian Adimora, Hongyang Sun 0001
|
||
TL;DR — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.
|
||
Why notable — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.
|
||
MIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie et al."><meta property="og:locale" content="en_us"><meta property="og:type" content="article"><meta property="article:section" content="cloud-edge"><meta property="article:published_time" content="2025-01-01T00:00:00+00:00"><meta property="article:modified_time" content="2025-01-01T00:00:00+00:00"><meta name=twitter:card content="summary"><meta name=twitter:title content="TPDS 2025 Digest"><meta name=twitter:description content="12 papers selected.
|
||
HARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management Kyrian Adimora, Hongyang Sun 0001
|
||
TL;DR — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.
|
||
Why notable — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.
|
||
MIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie et al."><script type=application/ld+json>{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Edge and Cloud Systems","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/"},{"@type":"ListItem","position":2,"name":"Digests","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/"},{"@type":"ListItem","position":3,"name":"TPDS 2025 Digest","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2025/"}]}</script><script type=application/ld+json>{"@context":"https://schema.org","@type":"BlogPosting","headline":"TPDS 2025 Digest","name":"TPDS 2025 Digest","description":"12 papers selected.\nHARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management Kyrian Adimora, Hongyang Sun 0001\nTL;DR — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.\nWhy notable — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.\nMIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie et al.\n","keywords":[],"articleBody":"12 papers selected.\nHARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management Kyrian Adimora, Hongyang Sun 0001\nTL;DR — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.\nWhy notable — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.\nMIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie et al.\nTL;DR — MIST reduces MPI job startup and termination latency to near-instant on the Tianhe supercomputer by redesigning the process-management and communication-bootstrap path.\nWhy notable — Startup overhead is a significant fraction of short-job turnaround time at scale; MIST’s results on a top-ranked system provide a concrete reference for HPC runtime developers.\nScheduling With Lightweight Predictions in Power-Constrained HPC Platforms Danilo Carastan-Santos, Georges Da Costa, Igor Fontana De Nardin, Millian Poquet et al.\nTL;DR — This paper develops a scheduling framework that uses lightweight runtime predictions to respect power caps on HPC systems while minimizing job slowdown.\nWhy notable — Power capping is now a first-class constraint on modern supercomputers, and this work from leading European HPC scheduling researchers offers practical, deployable algorithms.\nPipeMesh: Achieving Memory-Efficient Computation-Communication Overlap for Training Large Language Models Fanxin Li, Shixiong Zhao, Yuhao Qing, Jianyu Jiang et al.\nTL;DR — PipeMesh overlaps pipeline-parallel computation and communication for LLM training while carefully managing memory to avoid out-of-memory failures.\nWhy notable — Communication-computation overlap is one of the most impactful levers for LLM training efficiency, and PipeMesh’s memory-awareness addresses the key practical constraint.\nEfficientMoE: Optimizing Mixture-of-Experts Model Training With Adaptive Load Balance Yan Zeng, Chengchuang Huang, Yipeng Mei, Lifu Zhang 0004 et al.\nTL;DR — EfficientMoE introduces an adaptive load-balancing strategy for Mixture-of-Experts training that equalizes expert utilization and reduces communication bottlenecks.\nWhy notable — MoE models are central to frontier LLM architectures, and load imbalance is their primary training inefficiency; this work provides both analysis and a practical solution.\nSSpMM: Efficiently Scalable SpMM Kernels Across Multiple Generations of Tensor Cores Zeyu Xue, Mei Wen, Jianchao Yang, Minjin Tang et al.\nTL;DR — SSpMM delivers portable, high-performance sparse-matrix dense-matrix multiplication kernels that scale efficiently across Ampere, Hopper, and future Tensor Core generations.\nWhy notable — SpMM is a bottleneck in GNN training and scientific computing; cross-generation portability without performance loss is a significant contribution for the GPU computing community.\nIceFrog: A Layer-Elastic Scheduling System for Deep Learning Training in GPU Clusters Wei Gao 0064, Zhuoyuan Ouyang, Peng Sun 0006, Tianwei Zhang 0004 et al.\nTL;DR — IceFrog dynamically adjusts the number of pipeline stages (layers) assigned to each GPU during training to adapt to cluster heterogeneity and improve utilization.\nWhy notable — Layer elasticity is a novel dimension of flexibility in distributed DNN training; IceFrog’s scheduler provides measurable throughput gains in realistic heterogeneous GPU clusters.\nElastic Relaxation of Concurrent Data Structures Kåre von Geijer, Philippas Tsigas\nTL;DR — This paper introduces a formal framework and concrete algorithms for elastic relaxation of concurrent data structures, allowing tunable trade-offs between consistency and throughput.\nWhy notable — Tsigas’s group advances concurrent data-structure theory with a unifying formalism that subsumes many ad-hoc relaxed designs and enables provable guarantees.\nApproximation Algorithms for Scheduling With/Without Deadline Constraints Where Rejection Costs are Proportional to Processing Times Olivier Beaumont, Rémi Bouzel, Lionel Eyraud-Dubois, Esragul Korkmaz et al.\nTL;DR — This paper derives new approximation algorithms with tight ratios for online and offline scheduling problems where rejected jobs incur costs proportional to their processing times.\nWhy notable — The theoretical results close open gaps in parallel scheduling complexity and are directly applicable to cloud and HPC batch schedulers that must handle job rejection.\nEdgeHydra: Fault-Tolerant Edge Data Distribution Based on Erasure Coding Qiang He 0001, Guobiao Zhang, Jiawei Wang 0003, Ruikun Luo et al.\nTL;DR — EdgeHydra uses erasure coding tailored to edge-node failure patterns to provide fault-tolerant data distribution with low redundancy overhead at the network edge.\nWhy notable — Fault tolerance at the edge is an increasingly critical requirement, and this system’s erasure-coding approach significantly outperforms replication in storage efficiency.\nTwo-Dimensional Balanced Partitioning and Efficient Caching for Distributed Graph Analysis Shuai Lin, Rui Wang 0076, Yongkun Li 0001, Yinlong Xu 0001 et al.\nTL;DR — This paper proposes a 2D balanced graph partitioning scheme combined with a caching policy that jointly minimizes communication and replication costs in distributed graph systems.\nWhy notable — Graph partitioning and caching are co-dependent problems rarely treated together; the combined optimization yields substantial performance improvements with strong theoretical backing.\nTowards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms Zhongyi Lin, Ning Sun, Pallab Bhattacharya, Xizhou Feng et al.\nTL;DR — This work builds a platform-agnostic performance model for distributed ML training that accurately predicts training throughput across diverse multi-GPU configurations without per-system profiling.\nWhy notable — A universal modeling framework from Owens’s group removes the need for expensive empirical searches when tuning distributed training configurations, benefiting the entire ML systems community.\n","wordCount":"838","inLanguage":"en","datePublished":"2025-01-01T00:00:00Z","dateModified":"2025-01-01T00:00:00Z","author":{"@type":"Person","name":"Publish Assistant"},"mainEntityOfPage":{"@type":"WebPage","@id":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2025/"},"publisher":{"@type":"Organization","name":"Publish Assistant","logo":{"@type":"ImageObject","url":"https://pub.sqrt.fr/vincent/publish-assistant/favicon.ico"}}}</script></head><body id=top><header class=header><nav class=header-nav><div class=logo><a href=https://pub.sqrt.fr/vincent/publish-assistant/ accesskey=h title="Publish Assistant (Alt + H)">Publish Assistant</a>
|
||
<span class=logo-sep>/</span>
|
||
<a class=logo-topic href=/vincent/publish-assistant/cloud-edge/ title="Edge and Cloud Systems">Edge and Cloud Systems</a><div class=logo-switches><button id=theme-toggle class=theme-toggle accesskey=t title="(Alt + T)" aria-label="Toggle theme">
|
||
<svg class="moon" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1111.21 3 7 7 0 0021 12.79z"/></svg>
|
||
<svg class="sun" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg></button></div></div><ul id=menu class=menu><li><a href=/vincent/publish-assistant/cloud-edge/venues/ title=Venues><span>Venues</span></a></li><li><a href=/vincent/publish-assistant/cloud-edge/calendar/ title=Calendar><span>Calendar</span></a></li><li><a href=/vincent/publish-assistant/cloud-edge/digests/ title=Digests><span class=active>Digests</span></a></li></ul></nav></header><main class=main><article class=post-single><header class=post-header><nav class=breadcrumbs role=navigation aria-label=Breadcrumb><a href=/vincent/publish-assistant/cloud-edge/digests/>Digests</a>
|
||
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevron-right"><polyline points="9 18 15 12 9 6"/></svg></nav><h1 class="post-title entry-hint-parent">TPDS 2025 Digest</h1><div class=post-meta><span title='2025-01-01 00:00:00 +0000 UTC'>January 1, 2025</span> · <span>Publish Assistant</span></div></header><div class="post-content md-content"><p>12 papers selected.</p><hr><h3 id=harmonic-uncertainty-aware-multi-objective-optimization-for-energy-efficient-hpc-resource-management>HARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management<a hidden class=anchor aria-hidden=true href=#harmonic-uncertainty-aware-multi-objective-optimization-for-energy-efficient-hpc-resource-management>#</a></h3><p><em>Kyrian Adimora, Hongyang Sun 0001</em></p><p><strong>TL;DR</strong> — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation.</p><p><strong>Why notable</strong> — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers.</p><hr><h3 id=mist-towards-mpi-instant-startup-and-termination-on-tianhe-hpc-systems>MIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems<a hidden class=anchor aria-hidden=true href=#mist-towards-mpi-instant-startup-and-termination-on-tianhe-hpc-systems>#</a></h3><p><em>Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie <em>et al.</em></em></p><p><strong>TL;DR</strong> — MIST reduces MPI job startup and termination latency to near-instant on the Tianhe supercomputer by redesigning the process-management and communication-bootstrap path.</p><p><strong>Why notable</strong> — Startup overhead is a significant fraction of short-job turnaround time at scale; MIST’s results on a top-ranked system provide a concrete reference for HPC runtime developers.</p><hr><h3 id=scheduling-with-lightweight-predictions-in-power-constrained-hpc-platforms>Scheduling With Lightweight Predictions in Power-Constrained HPC Platforms<a hidden class=anchor aria-hidden=true href=#scheduling-with-lightweight-predictions-in-power-constrained-hpc-platforms>#</a></h3><p><em>Danilo Carastan-Santos, Georges Da Costa, Igor Fontana De Nardin, Millian Poquet <em>et al.</em></em></p><p><strong>TL;DR</strong> — This paper develops a scheduling framework that uses lightweight runtime predictions to respect power caps on HPC systems while minimizing job slowdown.</p><p><strong>Why notable</strong> — Power capping is now a first-class constraint on modern supercomputers, and this work from leading European HPC scheduling researchers offers practical, deployable algorithms.</p><hr><h3 id=pipemesh-achieving-memory-efficient-computation-communication-overlap-for-training-large-language-models>PipeMesh: Achieving Memory-Efficient Computation-Communication Overlap for Training Large Language Models<a hidden class=anchor aria-hidden=true href=#pipemesh-achieving-memory-efficient-computation-communication-overlap-for-training-large-language-models>#</a></h3><p><em>Fanxin Li, Shixiong Zhao, Yuhao Qing, Jianyu Jiang <em>et al.</em></em></p><p><strong>TL;DR</strong> — PipeMesh overlaps pipeline-parallel computation and communication for LLM training while carefully managing memory to avoid out-of-memory failures.</p><p><strong>Why notable</strong> — Communication-computation overlap is one of the most impactful levers for LLM training efficiency, and PipeMesh’s memory-awareness addresses the key practical constraint.</p><hr><h3 id=efficientmoe-optimizing-mixture-of-experts-model-training-with-adaptive-load-balance>EfficientMoE: Optimizing Mixture-of-Experts Model Training With Adaptive Load Balance<a hidden class=anchor aria-hidden=true href=#efficientmoe-optimizing-mixture-of-experts-model-training-with-adaptive-load-balance>#</a></h3><p><em>Yan Zeng, Chengchuang Huang, Yipeng Mei, Lifu Zhang 0004 <em>et al.</em></em></p><p><strong>TL;DR</strong> — EfficientMoE introduces an adaptive load-balancing strategy for Mixture-of-Experts training that equalizes expert utilization and reduces communication bottlenecks.</p><p><strong>Why notable</strong> — MoE models are central to frontier LLM architectures, and load imbalance is their primary training inefficiency; this work provides both analysis and a practical solution.</p><hr><h3 id=sspmm-efficiently-scalable-spmm-kernels-across-multiple-generations-of-tensor-cores>SSpMM: Efficiently Scalable SpMM Kernels Across Multiple Generations of Tensor Cores<a hidden class=anchor aria-hidden=true href=#sspmm-efficiently-scalable-spmm-kernels-across-multiple-generations-of-tensor-cores>#</a></h3><p><em>Zeyu Xue, Mei Wen, Jianchao Yang, Minjin Tang <em>et al.</em></em></p><p><strong>TL;DR</strong> — SSpMM delivers portable, high-performance sparse-matrix dense-matrix multiplication kernels that scale efficiently across Ampere, Hopper, and future Tensor Core generations.</p><p><strong>Why notable</strong> — SpMM is a bottleneck in GNN training and scientific computing; cross-generation portability without performance loss is a significant contribution for the GPU computing community.</p><hr><h3 id=icefrog-a-layer-elastic-scheduling-system-for-deep-learning-training-in-gpu-clusters>IceFrog: A Layer-Elastic Scheduling System for Deep Learning Training in GPU Clusters<a hidden class=anchor aria-hidden=true href=#icefrog-a-layer-elastic-scheduling-system-for-deep-learning-training-in-gpu-clusters>#</a></h3><p><em>Wei Gao 0064, Zhuoyuan Ouyang, Peng Sun 0006, Tianwei Zhang 0004 <em>et al.</em></em></p><p><strong>TL;DR</strong> — IceFrog dynamically adjusts the number of pipeline stages (layers) assigned to each GPU during training to adapt to cluster heterogeneity and improve utilization.</p><p><strong>Why notable</strong> — Layer elasticity is a novel dimension of flexibility in distributed DNN training; IceFrog’s scheduler provides measurable throughput gains in realistic heterogeneous GPU clusters.</p><hr><h3 id=elastic-relaxation-of-concurrent-data-structures>Elastic Relaxation of Concurrent Data Structures<a hidden class=anchor aria-hidden=true href=#elastic-relaxation-of-concurrent-data-structures>#</a></h3><p><em>Kåre von Geijer, Philippas Tsigas</em></p><p><strong>TL;DR</strong> — This paper introduces a formal framework and concrete algorithms for elastic relaxation of concurrent data structures, allowing tunable trade-offs between consistency and throughput.</p><p><strong>Why notable</strong> — Tsigas’s group advances concurrent data-structure theory with a unifying formalism that subsumes many ad-hoc relaxed designs and enables provable guarantees.</p><hr><h3 id=approximation-algorithms-for-scheduling-withwithout-deadline-constraints-where-rejection-costs-are-proportional-to-processing-times>Approximation Algorithms for Scheduling With/Without Deadline Constraints Where Rejection Costs are Proportional to Processing Times<a hidden class=anchor aria-hidden=true href=#approximation-algorithms-for-scheduling-withwithout-deadline-constraints-where-rejection-costs-are-proportional-to-processing-times>#</a></h3><p><em>Olivier Beaumont, Rémi Bouzel, Lionel Eyraud-Dubois, Esragul Korkmaz <em>et al.</em></em></p><p><strong>TL;DR</strong> — This paper derives new approximation algorithms with tight ratios for online and offline scheduling problems where rejected jobs incur costs proportional to their processing times.</p><p><strong>Why notable</strong> — The theoretical results close open gaps in parallel scheduling complexity and are directly applicable to cloud and HPC batch schedulers that must handle job rejection.</p><hr><h3 id=edgehydra-fault-tolerant-edge-data-distribution-based-on-erasure-coding>EdgeHydra: Fault-Tolerant Edge Data Distribution Based on Erasure Coding<a hidden class=anchor aria-hidden=true href=#edgehydra-fault-tolerant-edge-data-distribution-based-on-erasure-coding>#</a></h3><p><em>Qiang He 0001, Guobiao Zhang, Jiawei Wang 0003, Ruikun Luo <em>et al.</em></em></p><p><strong>TL;DR</strong> — EdgeHydra uses erasure coding tailored to edge-node failure patterns to provide fault-tolerant data distribution with low redundancy overhead at the network edge.</p><p><strong>Why notable</strong> — Fault tolerance at the edge is an increasingly critical requirement, and this system’s erasure-coding approach significantly outperforms replication in storage efficiency.</p><hr><h3 id=two-dimensional-balanced-partitioning-and-efficient-caching-for-distributed-graph-analysis>Two-Dimensional Balanced Partitioning and Efficient Caching for Distributed Graph Analysis<a hidden class=anchor aria-hidden=true href=#two-dimensional-balanced-partitioning-and-efficient-caching-for-distributed-graph-analysis>#</a></h3><p><em>Shuai Lin, Rui Wang 0076, Yongkun Li 0001, Yinlong Xu 0001 <em>et al.</em></em></p><p><strong>TL;DR</strong> — This paper proposes a 2D balanced graph partitioning scheme combined with a caching policy that jointly minimizes communication and replication costs in distributed graph systems.</p><p><strong>Why notable</strong> — Graph partitioning and caching are co-dependent problems rarely treated together; the combined optimization yields substantial performance improvements with strong theoretical backing.</p><hr><h3 id=towards-universal-performance-modeling-for-machine-learning-training-on-multi-gpu-platforms>Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms<a hidden class=anchor aria-hidden=true href=#towards-universal-performance-modeling-for-machine-learning-training-on-multi-gpu-platforms>#</a></h3><p><em>Zhongyi Lin, Ning Sun, Pallab Bhattacharya, Xizhou Feng <em>et al.</em></em></p><p><strong>TL;DR</strong> — This work builds a platform-agnostic performance model for distributed ML training that accurately predicts training throughput across diverse multi-GPU configurations without per-system profiling.</p><p><strong>Why notable</strong> — A universal modeling framework from Owens’s group removes the need for expensive empirical searches when tuning distributed training configurations, benefiting the entire ML systems community.</p></div><footer class=post-footer><ul class=post-tags></ul><nav class=paginav><a class=prev href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tocs-2025/><span class=title>« Prev</span>
|
||
<span>TOCS 2025 Digest</span>
|
||
</a><a class=next href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/middleware-2024/><span class=title>Next »</span>
|
||
<span>Middleware 2024 Digest</span></a></nav></footer></article></main><footer class=footer><span>© 2026 <a href=https://pub.sqrt.fr/vincent/publish-assistant/>Publish Assistant</a></span> ·
|
||
<span>Powered by
|
||
<a href="https://gohugo.io/?utm_source=papermod" rel=noopener target=_blank>Hugo</a> &
|
||
<a href=https://github.com/adityatelange/hugo-PaperMod/ rel=noopener target=_blank>PaperMod</a></span></footer><a href=#top id=top-link class="top-link hidden" aria-label="go to top" title="Go to Top (Alt + G)" accesskey=g><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevrons-up"><polyline points="17 11 12 6 7 11"/><polyline points="17 18 12 13 7 18"/></svg>
|
||
</a><script>let menu=document.getElementById("menu");if(menu){const e=localStorage.getItem("menu-scroll-position");e&&(menu.scrollLeft=parseInt(e,10)),menu.onscroll=function(){localStorage.setItem("menu-scroll-position",menu.scrollLeft)}}document.querySelectorAll('a[href^="#"]').forEach(e=>{e.addEventListener("click",function(e){e.preventDefault();var t=this.getAttribute("href").substr(1);window.matchMedia("(prefers-reduced-motion: reduce)").matches?document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView():document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView({behavior:"smooth"}),t==="top"?history.replaceState(null,null," "):history.pushState(null,null,`#${t}`)})})</script><script>var toplink=document.getElementById("top-link");window.onscroll=function(){const e=window.innerHeight;document.body.scrollTop>e||document.documentElement.scrollTop>e?toplink.classList.remove("hidden"):toplink.classList.add("hidden")}</script><script>document.getElementById("theme-toggle").addEventListener("click",()=>{const e=document.querySelector("html");e.dataset.theme==="dark"?(e.dataset.theme="light",localStorage.setItem("pref-theme","light")):(e.dataset.theme="dark",localStorage.setItem("pref-theme","dark"))})</script></body></html> |