Files
2026-08-18 13:39:21 +00:00

27 lines
26 KiB
HTML
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!doctype html><html lang=en dir=auto data-theme=auto><head><meta charset=utf-8><meta http-equiv=X-UA-Compatible content="IE=edge"><meta name=viewport content="width=device-width,initial-scale=1,shrink-to-fit=no"><meta name=robots content="index, follow"><title>TPDS 2024 Digest | Publish Assistant</title><meta name=keywords content><meta name=description content="12 papers selected.
Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning
Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.
TL;DR — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.
Why notable — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.
AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost
Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan et al."><meta name=author content="Publish Assistant"><link rel=canonical href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2024/><link crossorigin=anonymous href=/vincent/publish-assistant/assets/css/stylesheet.d72f07832e13c592b3edba91680bfe70f01daac396179bcace0ac36e8e0494c6.css integrity="sha256-1y8Hgy4TxZKz7bqRaAv+cPAdqsOWF5vKzgrDbo4ElMY=" rel="preload stylesheet" as=style><link rel=icon href=https://pub.sqrt.fr/vincent/publish-assistant/favicon.ico><link rel=icon type=image/png sizes=16x16 href=https://pub.sqrt.fr/vincent/publish-assistant/favicon-16x16.png><link rel=icon type=image/png sizes=32x32 href=https://pub.sqrt.fr/vincent/publish-assistant/favicon-32x32.png><link rel=apple-touch-icon href=https://pub.sqrt.fr/vincent/publish-assistant/apple-touch-icon.png><link rel=mask-icon href=https://pub.sqrt.fr/vincent/publish-assistant/safari-pinned-tab.svg><meta name=theme-color content="#2e2e33"><meta name=msapplication-TileColor content="#2e2e33"><link rel=alternate hreflang=en href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2024/><noscript><style>#theme-toggle,.top-link{display:none}</style><style>@media(prefers-color-scheme:dark){:root{--theme:rgb(29, 30, 32);--entry:rgb(46, 46, 51);--primary:rgb(218, 218, 219);--secondary:rgb(155, 156, 157);--tertiary:rgb(65, 66, 68);--content:rgb(196, 196, 197);--code-block-bg:rgb(46, 46, 51);--code-bg:rgb(55, 56, 62);--border:rgb(51, 51, 51);color-scheme:dark}.list{background:var(--theme)}.toc{background:var(--entry)}}</style></noscript><script>localStorage.getItem("pref-theme")==="dark"?document.querySelector("html").dataset.theme="dark":localStorage.getItem("pref-theme")==="light"?document.querySelector("html").dataset.theme="light":window.matchMedia("(prefers-color-scheme: dark)").matches?document.querySelector("html").dataset.theme="dark":document.querySelector("html").dataset.theme="light"</script><meta property="og:url" content="https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2024/"><meta property="og:site_name" content="Publish Assistant"><meta property="og:title" content="TPDS 2024 Digest"><meta property="og:description" content="12 papers selected.
Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.
TL;DR — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.
Why notable — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.
AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan et al."><meta property="og:locale" content="en_us"><meta property="og:type" content="article"><meta property="article:section" content="cloud-edge"><meta property="article:published_time" content="2024-01-01T00:00:00+00:00"><meta property="article:modified_time" content="2024-01-01T00:00:00+00:00"><meta name=twitter:card content="summary"><meta name=twitter:title content="TPDS 2024 Digest"><meta name=twitter:description content="12 papers selected.
Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.
TL;DR — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.
Why notable — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.
AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan et al."><script type=application/ld+json>{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Edge and Cloud Systems","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/"},{"@type":"ListItem","position":2,"name":"Digests","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/"},{"@type":"ListItem","position":3,"name":"TPDS 2024 Digest","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2024/"}]}</script><script type=application/ld+json>{"@context":"https://schema.org","@type":"BlogPosting","headline":"TPDS 2024 Digest","name":"TPDS 2024 Digest","description":"12 papers selected.\nRuntime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.\nTL;DR — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.\nWhy notable — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.\nAutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan et al.\n","keywords":[],"articleBody":"12 papers selected.\nRuntime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.\nTL;DR — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.\nWhy notable — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.\nAutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan et al.\nTL;DR — AutoDDL automatically searches for the distributed DNN training strategy that minimizes communication bandwidth cost while meeting performance targets.\nWhy notable — Combining Hoeflers communication-model expertise with automatic strategy search, this paper is essential reading for practitioners scaling DNN training across large clusters.\nPeakFS: An Ultra-High Performance Parallel File System via Computing-Network-Storage Co-Optimization for HPC Applications Yixiao Chen, Haomai Yang, Kai Lu 0002, Wenlve Huang et al.\nTL;DR — PeakFS co-optimizes compute, network, and storage layers of a parallel file system to deliver ultra-high I/O throughput for HPC workloads.\nWhy notable — Its holistic co-design perspective sets a new performance baseline for HPC storage and provides actionable insights for next-generation parallel file system architects.\nFormal Definitions and Performance Comparison of Consistency Models for Parallel File Systems Chen Wang 0004, Kathryn M. Mohror, Marc Snir\nTL;DR — This paper formalizes consistency models used by parallel file systems and provides the first systematic empirical comparison of their performance trade-offs.\nWhy notable — Rigorous formal treatment from Snir and Mohror clarifies long-standing ambiguities in HPC storage semantics, making it an important reference for storage system designers.\nMalleability in Modern HPC Systems: Current Experiences, Challenges, and Future Opportunities Ahmad Tarraf, Martin Schreiber 0001, Alberto Cascajo, Jean-Baptiste Besnard et al.\nTL;DR — A comprehensive survey of dynamic resource malleability in HPC, covering runtime systems, job schedulers, and application-level support with lessons from production systems.\nWhy notable — As energy-aware and burst-resilient HPC scheduling becomes critical, this broad community-driven survey is the definitive starting point for research on malleable HPC runtimes.\nPyxis: Scheduling Mixed Tasks in Disaggregated Datacenters Sheng Qi, Chao Jin, Mosharaf Chowdhury, Zhenming Liu et al.\nTL;DR — Pyxis is a scheduler for disaggregated datacenters that jointly manages latency-sensitive and batch tasks by exploiting flexible resource pooling across the disaggregated fabric.\nWhy notable — It tackles one of the central open problems in cloud scheduling—multi-tenancy under disaggregation—with rigorous analysis and demonstrated gains on real workloads.\nSwift: Expedited Failure Recovery for Large-Scale DNN Training Yuchen Zhong, Guangming Sheng, Juncheng Liu, Jinhui Yuan et al.\nTL;DR — Swift dramatically reduces checkpoint and recovery overhead for large-scale DNN training by combining lightweight in-memory snapshots with selective recomputation.\nWhy notable — As training runs on hundreds of GPUs grow longer and failures become inevitable, Swifts fault-tolerance approach directly addresses a practical bottleneck in modern deep-learning infrastructure.\nFastLoad: Speeding Up Data Loading of Both Sparse Matrix and Vector for SpMV on GPUs Jinyu Hu, Huizhang Luo, Hong Jiang 0001, Guoqing Xiao 0001 et al.\nTL;DR — FastLoad optimizes the memory-access pattern for loading both the sparse matrix and the dense vector in SpMV on GPUs, yielding significant throughput improvements.\nWhy notable — SpMV is a foundational kernel for scientific computing and graph analytics; this papers memory-access analysis and optimizations benefit a wide class of GPU applications.\nKLNK: Expanding Page Boundaries in a Distributed Shared Memory System Yiwei Ci, Michael R. Lyu, Zhan Zhang 0002, De-Cheng Zuo et al.\nTL;DR — KLNK extends distributed shared memory page granularity to reduce false sharing and improve throughput for irregular access patterns.\nWhy notable — It addresses a classic but unsolved bottleneck in DSM systems with a practical, page-table-level mechanism applicable to emerging disaggregated memory architectures.\nEnabling Efficient Erasure Coding in Disaggregated Memory Systems Qiliang Li, Liangliang Xu, Yongkun Li 0001, Min Lyu et al.\nTL;DR — This paper designs an erasure-coding scheme tailored to disaggregated memory, exploiting its unique bandwidth topology to achieve fault tolerance with low overhead.\nWhy notable — Fault tolerance in disaggregated memory is an open problem of growing importance; this work provides concrete mechanisms and strong performance results.\nSimple, Fast and Widely Applicable Concurrent Memory Reclamation via Neutralization Ajay Singh 0002, Trevor Alexander Brown, Ali José Mashtizadeh\nTL;DR — Neutralization is a new mechanism for safe memory reclamation in lock-free data structures that is simpler, faster, and more portable than prior approaches.\nWhy notable — Safe memory reclamation is a pervasive challenge in concurrent programming; this algorithms breadth of applicability and performance improvements make it highly reusable.\nDeepTM: Efficient Tensor Management in Heterogeneous Memory for DNN Training Haoran Zhou, Wei Rang, Hongyang Chen 0001, Xiaobo Zhou 0002 et al.\nTL;DR — DeepTM dynamically manages tensor placement across DRAM and NVM during DNN training to reduce memory pressure and improve throughput.\nWhy notable — As model sizes outpace GPU memory, heterogeneous memory management becomes critical; DeepTM provides a practical, training-aware solution with measurable benefits.\n","wordCount":"823","inLanguage":"en","datePublished":"2024-01-01T00:00:00Z","dateModified":"2024-01-01T00:00:00Z","author":{"@type":"Person","name":"Publish Assistant"},"mainEntityOfPage":{"@type":"WebPage","@id":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2024/"},"publisher":{"@type":"Organization","name":"Publish Assistant","logo":{"@type":"ImageObject","url":"https://pub.sqrt.fr/vincent/publish-assistant/favicon.ico"}}}</script></head><body id=top><header class=header><nav class=header-nav><div class=logo><a href=https://pub.sqrt.fr/vincent/publish-assistant/ accesskey=h title="Publish Assistant (Alt + H)">Publish Assistant</a>
<span class=logo-sep>/</span>
<a class=logo-topic href=/vincent/publish-assistant/cloud-edge/ title="Edge and Cloud Systems">Edge and Cloud Systems</a><div class=logo-switches><button id=theme-toggle class=theme-toggle accesskey=t title="(Alt + T)" aria-label="Toggle theme">
<svg class="moon" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1111.21 3 7 7 0 0021 12.79z"/></svg>
<svg class="sun" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg></button></div></div><ul id=menu class=menu><li><a href=/vincent/publish-assistant/cloud-edge/venues/ title=Venues><span>Venues</span></a></li><li><a href=/vincent/publish-assistant/cloud-edge/calendar/ title=Calendar><span>Calendar</span></a></li><li><a href=/vincent/publish-assistant/cloud-edge/digests/ title=Digests><span class=active>Digests</span></a></li></ul></nav></header><main class=main><article class=post-single><header class=post-header><nav class=breadcrumbs role=navigation aria-label=Breadcrumb><a href=/vincent/publish-assistant/cloud-edge/digests/>Digests</a>
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevron-right"><polyline points="9 18 15 12 9 6"/></svg></nav><h1 class="post-title entry-hint-parent">TPDS 2024 Digest</h1><div class=post-meta><span title='2024-01-01 00:00:00 +0000 UTC'>January 1, 2024</span>&nbsp;·&nbsp;<span>Publish Assistant</span></div></header><div class="post-content md-content"><p>12 papers selected.</p><hr><h3 id=runtime-performance-anomaly-diagnosis-in-production-hpc-systems-using-active-learning>Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning<a hidden class=anchor aria-hidden=true href=#runtime-performance-anomaly-diagnosis-in-production-hpc-systems-using-active-learning>#</a></h3><p><em>Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz <em>et al.</em></em></p><p><strong>TL;DR</strong> — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.</p><p><strong>Why notable</strong> — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.</p><hr><h3 id=autoddl-automatic-distributed-deep-learning-with-near-optimal-bandwidth-cost>AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost<a hidden class=anchor aria-hidden=true href=#autoddl-automatic-distributed-deep-learning-with-near-optimal-bandwidth-cost>#</a></h3><p><em>Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan <em>et al.</em></em></p><p><strong>TL;DR</strong> — AutoDDL automatically searches for the distributed DNN training strategy that minimizes communication bandwidth cost while meeting performance targets.</p><p><strong>Why notable</strong> — Combining Hoefler&rsquo;s communication-model expertise with automatic strategy search, this paper is essential reading for practitioners scaling DNN training across large clusters.</p><hr><h3 id=peakfs-an-ultra-high-performance-parallel-file-system-via-computing-network-storage-co-optimization-for-hpc-applications>PeakFS: An Ultra-High Performance Parallel File System via Computing-Network-Storage Co-Optimization for HPC Applications<a hidden class=anchor aria-hidden=true href=#peakfs-an-ultra-high-performance-parallel-file-system-via-computing-network-storage-co-optimization-for-hpc-applications>#</a></h3><p><em>Yixiao Chen, Haomai Yang, Kai Lu 0002, Wenlve Huang <em>et al.</em></em></p><p><strong>TL;DR</strong> — PeakFS co-optimizes compute, network, and storage layers of a parallel file system to deliver ultra-high I/O throughput for HPC workloads.</p><p><strong>Why notable</strong> — Its holistic co-design perspective sets a new performance baseline for HPC storage and provides actionable insights for next-generation parallel file system architects.</p><hr><h3 id=formal-definitions-and-performance-comparison-of-consistency-models-for-parallel-file-systems>Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems<a hidden class=anchor aria-hidden=true href=#formal-definitions-and-performance-comparison-of-consistency-models-for-parallel-file-systems>#</a></h3><p><em>Chen Wang 0004, Kathryn M. Mohror, Marc Snir</em></p><p><strong>TL;DR</strong> — This paper formalizes consistency models used by parallel file systems and provides the first systematic empirical comparison of their performance trade-offs.</p><p><strong>Why notable</strong> — Rigorous formal treatment from Snir and Mohror clarifies long-standing ambiguities in HPC storage semantics, making it an important reference for storage system designers.</p><hr><h3 id=malleability-in-modern-hpc-systems-current-experiences-challenges-and-future-opportunities>Malleability in Modern HPC Systems: Current Experiences, Challenges, and Future Opportunities<a hidden class=anchor aria-hidden=true href=#malleability-in-modern-hpc-systems-current-experiences-challenges-and-future-opportunities>#</a></h3><p><em>Ahmad Tarraf, Martin Schreiber 0001, Alberto Cascajo, Jean-Baptiste Besnard <em>et al.</em></em></p><p><strong>TL;DR</strong> — A comprehensive survey of dynamic resource malleability in HPC, covering runtime systems, job schedulers, and application-level support with lessons from production systems.</p><p><strong>Why notable</strong> — As energy-aware and burst-resilient HPC scheduling becomes critical, this broad community-driven survey is the definitive starting point for research on malleable HPC runtimes.</p><hr><h3 id=pyxis-scheduling-mixed-tasks-in-disaggregated-datacenters>Pyxis: Scheduling Mixed Tasks in Disaggregated Datacenters<a hidden class=anchor aria-hidden=true href=#pyxis-scheduling-mixed-tasks-in-disaggregated-datacenters>#</a></h3><p><em>Sheng Qi, Chao Jin, Mosharaf Chowdhury, Zhenming Liu <em>et al.</em></em></p><p><strong>TL;DR</strong> — Pyxis is a scheduler for disaggregated datacenters that jointly manages latency-sensitive and batch tasks by exploiting flexible resource pooling across the disaggregated fabric.</p><p><strong>Why notable</strong> — It tackles one of the central open problems in cloud scheduling—multi-tenancy under disaggregation—with rigorous analysis and demonstrated gains on real workloads.</p><hr><h3 id=swift-expedited-failure-recovery-for-large-scale-dnn-training>Swift: Expedited Failure Recovery for Large-Scale DNN Training<a hidden class=anchor aria-hidden=true href=#swift-expedited-failure-recovery-for-large-scale-dnn-training>#</a></h3><p><em>Yuchen Zhong, Guangming Sheng, Juncheng Liu, Jinhui Yuan <em>et al.</em></em></p><p><strong>TL;DR</strong> — Swift dramatically reduces checkpoint and recovery overhead for large-scale DNN training by combining lightweight in-memory snapshots with selective recomputation.</p><p><strong>Why notable</strong> — As training runs on hundreds of GPUs grow longer and failures become inevitable, Swift&rsquo;s fault-tolerance approach directly addresses a practical bottleneck in modern deep-learning infrastructure.</p><hr><h3 id=fastload-speeding-up-data-loading-of-both-sparse-matrix-and-vector-for-spmv-on-gpus>FastLoad: Speeding Up Data Loading of Both Sparse Matrix and Vector for SpMV on GPUs<a hidden class=anchor aria-hidden=true href=#fastload-speeding-up-data-loading-of-both-sparse-matrix-and-vector-for-spmv-on-gpus>#</a></h3><p><em>Jinyu Hu, Huizhang Luo, Hong Jiang 0001, Guoqing Xiao 0001 <em>et al.</em></em></p><p><strong>TL;DR</strong> — FastLoad optimizes the memory-access pattern for loading both the sparse matrix and the dense vector in SpMV on GPUs, yielding significant throughput improvements.</p><p><strong>Why notable</strong> — SpMV is a foundational kernel for scientific computing and graph analytics; this paper&rsquo;s memory-access analysis and optimizations benefit a wide class of GPU applications.</p><hr><h3 id=klnk-expanding-page-boundaries-in-a-distributed-shared-memory-system>KLNK: Expanding Page Boundaries in a Distributed Shared Memory System<a hidden class=anchor aria-hidden=true href=#klnk-expanding-page-boundaries-in-a-distributed-shared-memory-system>#</a></h3><p><em>Yiwei Ci, Michael R. Lyu, Zhan Zhang 0002, De-Cheng Zuo <em>et al.</em></em></p><p><strong>TL;DR</strong> — KLNK extends distributed shared memory page granularity to reduce false sharing and improve throughput for irregular access patterns.</p><p><strong>Why notable</strong> — It addresses a classic but unsolved bottleneck in DSM systems with a practical, page-table-level mechanism applicable to emerging disaggregated memory architectures.</p><hr><h3 id=enabling-efficient-erasure-coding-in-disaggregated-memory-systems>Enabling Efficient Erasure Coding in Disaggregated Memory Systems<a hidden class=anchor aria-hidden=true href=#enabling-efficient-erasure-coding-in-disaggregated-memory-systems>#</a></h3><p><em>Qiliang Li, Liangliang Xu, Yongkun Li 0001, Min Lyu <em>et al.</em></em></p><p><strong>TL;DR</strong> — This paper designs an erasure-coding scheme tailored to disaggregated memory, exploiting its unique bandwidth topology to achieve fault tolerance with low overhead.</p><p><strong>Why notable</strong> — Fault tolerance in disaggregated memory is an open problem of growing importance; this work provides concrete mechanisms and strong performance results.</p><hr><h3 id=simple-fast-and-widely-applicable-concurrent-memory-reclamation-via-neutralization>Simple, Fast and Widely Applicable Concurrent Memory Reclamation via Neutralization<a hidden class=anchor aria-hidden=true href=#simple-fast-and-widely-applicable-concurrent-memory-reclamation-via-neutralization>#</a></h3><p><em>Ajay Singh 0002, Trevor Alexander Brown, Ali José Mashtizadeh</em></p><p><strong>TL;DR</strong> — Neutralization is a new mechanism for safe memory reclamation in lock-free data structures that is simpler, faster, and more portable than prior approaches.</p><p><strong>Why notable</strong> — Safe memory reclamation is a pervasive challenge in concurrent programming; this algorithm&rsquo;s breadth of applicability and performance improvements make it highly reusable.</p><hr><h3 id=deeptm-efficient-tensor-management-in-heterogeneous-memory-for-dnn-training>DeepTM: Efficient Tensor Management in Heterogeneous Memory for DNN Training<a hidden class=anchor aria-hidden=true href=#deeptm-efficient-tensor-management-in-heterogeneous-memory-for-dnn-training>#</a></h3><p><em>Haoran Zhou, Wei Rang, Hongyang Chen 0001, Xiaobo Zhou 0002 <em>et al.</em></em></p><p><strong>TL;DR</strong> — DeepTM dynamically manages tensor placement across DRAM and NVM during DNN training to reduce memory pressure and improve throughput.</p><p><strong>Why notable</strong> — As model sizes outpace GPU memory, heterogeneous memory management becomes critical; DeepTM provides a practical, training-aware solution with measurable benefits.</p></div><footer class=post-footer><ul class=post-tags></ul><nav class=paginav><a class=prev href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tocs-2024/><span class=title>« Prev</span>
<span>TOCS 2024 Digest</span></a></nav></footer></article></main><footer class=footer><span>&copy; 2026 <a href=https://pub.sqrt.fr/vincent/publish-assistant/>Publish Assistant</a></span> ·
<span>Powered by
<a href="https://gohugo.io/?utm_source=papermod" rel=noopener target=_blank>Hugo</a> &
<a href=https://github.com/adityatelange/hugo-PaperMod/ rel=noopener target=_blank>PaperMod</a></span></footer><a href=#top id=top-link class="top-link hidden" aria-label="go to top" title="Go to Top (Alt + G)" accesskey=g><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevrons-up"><polyline points="17 11 12 6 7 11"/><polyline points="17 18 12 13 7 18"/></svg>
</a><script>let menu=document.getElementById("menu");if(menu){const e=localStorage.getItem("menu-scroll-position");e&&(menu.scrollLeft=parseInt(e,10)),menu.onscroll=function(){localStorage.setItem("menu-scroll-position",menu.scrollLeft)}}document.querySelectorAll('a[href^="#"]').forEach(e=>{e.addEventListener("click",function(e){e.preventDefault();var t=this.getAttribute("href").substr(1);window.matchMedia("(prefers-reduced-motion: reduce)").matches?document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView():document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView({behavior:"smooth"}),t==="top"?history.replaceState(null,null," "):history.pushState(null,null,`#${t}`)})})</script><script>var toplink=document.getElementById("top-link");window.onscroll=function(){const e=window.innerHeight;document.body.scrollTop>e||document.documentElement.scrollTop>e?toplink.classList.remove("hidden"):toplink.classList.add("hidden")}</script><script>document.getElementById("theme-toggle").addEventListener("click",()=>{const e=document.querySelector("html");e.dataset.theme==="dark"?(e.dataset.theme="light",localStorage.setItem("pref-theme","light")):(e.dataset.theme="dark",localStorage.setItem("pref-theme","dark"))})</script></body></html>