27 lines
26 KiB
HTML
27 lines
26 KiB
HTML
<!doctype html><html lang=en dir=auto data-theme=auto><head><meta charset=utf-8><meta http-equiv=X-UA-Compatible content="IE=edge"><meta name=viewport content="width=device-width,initial-scale=1,shrink-to-fit=no"><meta name=robots content="index, follow"><title>TPDS 2024 Digest | Publish Assistant</title><meta name=keywords content><meta name=description content="12 papers selected.
|
||
|
||
Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning
|
||
Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.
|
||
TL;DR — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.
|
||
Why notable — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.
|
||
|
||
AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost
|
||
Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan et al."><meta name=author content="Publish Assistant"><link rel=canonical href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2024/><link crossorigin=anonymous href=/vincent/publish-assistant/assets/css/stylesheet.d72f07832e13c592b3edba91680bfe70f01daac396179bcace0ac36e8e0494c6.css integrity="sha256-1y8Hgy4TxZKz7bqRaAv+cPAdqsOWF5vKzgrDbo4ElMY=" rel="preload stylesheet" as=style><link rel=icon href=https://pub.sqrt.fr/vincent/publish-assistant/favicon.ico><link rel=icon type=image/png sizes=16x16 href=https://pub.sqrt.fr/vincent/publish-assistant/favicon-16x16.png><link rel=icon type=image/png sizes=32x32 href=https://pub.sqrt.fr/vincent/publish-assistant/favicon-32x32.png><link rel=apple-touch-icon href=https://pub.sqrt.fr/vincent/publish-assistant/apple-touch-icon.png><link rel=mask-icon href=https://pub.sqrt.fr/vincent/publish-assistant/safari-pinned-tab.svg><meta name=theme-color content="#2e2e33"><meta name=msapplication-TileColor content="#2e2e33"><link rel=alternate hreflang=en href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2024/><noscript><style>#theme-toggle,.top-link{display:none}</style><style>@media(prefers-color-scheme:dark){:root{--theme:rgb(29, 30, 32);--entry:rgb(46, 46, 51);--primary:rgb(218, 218, 219);--secondary:rgb(155, 156, 157);--tertiary:rgb(65, 66, 68);--content:rgb(196, 196, 197);--code-block-bg:rgb(46, 46, 51);--code-bg:rgb(55, 56, 62);--border:rgb(51, 51, 51);color-scheme:dark}.list{background:var(--theme)}.toc{background:var(--entry)}}</style></noscript><script>localStorage.getItem("pref-theme")==="dark"?document.querySelector("html").dataset.theme="dark":localStorage.getItem("pref-theme")==="light"?document.querySelector("html").dataset.theme="light":window.matchMedia("(prefers-color-scheme: dark)").matches?document.querySelector("html").dataset.theme="dark":document.querySelector("html").dataset.theme="light"</script><meta property="og:url" content="https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2024/"><meta property="og:site_name" content="Publish Assistant"><meta property="og:title" content="TPDS 2024 Digest"><meta property="og:description" content="12 papers selected.
|
||
Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.
|
||
TL;DR — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.
|
||
Why notable — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.
|
||
AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan et al."><meta property="og:locale" content="en_us"><meta property="og:type" content="article"><meta property="article:section" content="cloud-edge"><meta property="article:published_time" content="2024-01-01T00:00:00+00:00"><meta property="article:modified_time" content="2024-01-01T00:00:00+00:00"><meta name=twitter:card content="summary"><meta name=twitter:title content="TPDS 2024 Digest"><meta name=twitter:description content="12 papers selected.
|
||
Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.
|
||
TL;DR — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.
|
||
Why notable — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.
|
||
AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan et al."><script type=application/ld+json>{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Edge and Cloud Systems","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/"},{"@type":"ListItem","position":2,"name":"Digests","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/"},{"@type":"ListItem","position":3,"name":"TPDS 2024 Digest","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2024/"}]}</script><script type=application/ld+json>{"@context":"https://schema.org","@type":"BlogPosting","headline":"TPDS 2024 Digest","name":"TPDS 2024 Digest","description":"12 papers selected.\nRuntime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.\nTL;DR — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.\nWhy notable — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.\nAutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan et al.\n","keywords":[],"articleBody":"12 papers selected.\nRuntime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.\nTL;DR — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.\nWhy notable — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.\nAutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan et al.\nTL;DR — AutoDDL automatically searches for the distributed DNN training strategy that minimizes communication bandwidth cost while meeting performance targets.\nWhy notable — Combining Hoefler’s communication-model expertise with automatic strategy search, this paper is essential reading for practitioners scaling DNN training across large clusters.\nPeakFS: An Ultra-High Performance Parallel File System via Computing-Network-Storage Co-Optimization for HPC Applications Yixiao Chen, Haomai Yang, Kai Lu 0002, Wenlve Huang et al.\nTL;DR — PeakFS co-optimizes compute, network, and storage layers of a parallel file system to deliver ultra-high I/O throughput for HPC workloads.\nWhy notable — Its holistic co-design perspective sets a new performance baseline for HPC storage and provides actionable insights for next-generation parallel file system architects.\nFormal Definitions and Performance Comparison of Consistency Models for Parallel File Systems Chen Wang 0004, Kathryn M. Mohror, Marc Snir\nTL;DR — This paper formalizes consistency models used by parallel file systems and provides the first systematic empirical comparison of their performance trade-offs.\nWhy notable — Rigorous formal treatment from Snir and Mohror clarifies long-standing ambiguities in HPC storage semantics, making it an important reference for storage system designers.\nMalleability in Modern HPC Systems: Current Experiences, Challenges, and Future Opportunities Ahmad Tarraf, Martin Schreiber 0001, Alberto Cascajo, Jean-Baptiste Besnard et al.\nTL;DR — A comprehensive survey of dynamic resource malleability in HPC, covering runtime systems, job schedulers, and application-level support with lessons from production systems.\nWhy notable — As energy-aware and burst-resilient HPC scheduling becomes critical, this broad community-driven survey is the definitive starting point for research on malleable HPC runtimes.\nPyxis: Scheduling Mixed Tasks in Disaggregated Datacenters Sheng Qi, Chao Jin, Mosharaf Chowdhury, Zhenming Liu et al.\nTL;DR — Pyxis is a scheduler for disaggregated datacenters that jointly manages latency-sensitive and batch tasks by exploiting flexible resource pooling across the disaggregated fabric.\nWhy notable — It tackles one of the central open problems in cloud scheduling—multi-tenancy under disaggregation—with rigorous analysis and demonstrated gains on real workloads.\nSwift: Expedited Failure Recovery for Large-Scale DNN Training Yuchen Zhong, Guangming Sheng, Juncheng Liu, Jinhui Yuan et al.\nTL;DR — Swift dramatically reduces checkpoint and recovery overhead for large-scale DNN training by combining lightweight in-memory snapshots with selective recomputation.\nWhy notable — As training runs on hundreds of GPUs grow longer and failures become inevitable, Swift’s fault-tolerance approach directly addresses a practical bottleneck in modern deep-learning infrastructure.\nFastLoad: Speeding Up Data Loading of Both Sparse Matrix and Vector for SpMV on GPUs Jinyu Hu, Huizhang Luo, Hong Jiang 0001, Guoqing Xiao 0001 et al.\nTL;DR — FastLoad optimizes the memory-access pattern for loading both the sparse matrix and the dense vector in SpMV on GPUs, yielding significant throughput improvements.\nWhy notable — SpMV is a foundational kernel for scientific computing and graph analytics; this paper’s memory-access analysis and optimizations benefit a wide class of GPU applications.\nKLNK: Expanding Page Boundaries in a Distributed Shared Memory System Yiwei Ci, Michael R. Lyu, Zhan Zhang 0002, De-Cheng Zuo et al.\nTL;DR — KLNK extends distributed shared memory page granularity to reduce false sharing and improve throughput for irregular access patterns.\nWhy notable — It addresses a classic but unsolved bottleneck in DSM systems with a practical, page-table-level mechanism applicable to emerging disaggregated memory architectures.\nEnabling Efficient Erasure Coding in Disaggregated Memory Systems Qiliang Li, Liangliang Xu, Yongkun Li 0001, Min Lyu et al.\nTL;DR — This paper designs an erasure-coding scheme tailored to disaggregated memory, exploiting its unique bandwidth topology to achieve fault tolerance with low overhead.\nWhy notable — Fault tolerance in disaggregated memory is an open problem of growing importance; this work provides concrete mechanisms and strong performance results.\nSimple, Fast and Widely Applicable Concurrent Memory Reclamation via Neutralization Ajay Singh 0002, Trevor Alexander Brown, Ali José Mashtizadeh\nTL;DR — Neutralization is a new mechanism for safe memory reclamation in lock-free data structures that is simpler, faster, and more portable than prior approaches.\nWhy notable — Safe memory reclamation is a pervasive challenge in concurrent programming; this algorithm’s breadth of applicability and performance improvements make it highly reusable.\nDeepTM: Efficient Tensor Management in Heterogeneous Memory for DNN Training Haoran Zhou, Wei Rang, Hongyang Chen 0001, Xiaobo Zhou 0002 et al.\nTL;DR — DeepTM dynamically manages tensor placement across DRAM and NVM during DNN training to reduce memory pressure and improve throughput.\nWhy notable — As model sizes outpace GPU memory, heterogeneous memory management becomes critical; DeepTM provides a practical, training-aware solution with measurable benefits.\n","wordCount":"823","inLanguage":"en","datePublished":"2024-01-01T00:00:00Z","dateModified":"2024-01-01T00:00:00Z","author":{"@type":"Person","name":"Publish Assistant"},"mainEntityOfPage":{"@type":"WebPage","@id":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tpds-2024/"},"publisher":{"@type":"Organization","name":"Publish Assistant","logo":{"@type":"ImageObject","url":"https://pub.sqrt.fr/vincent/publish-assistant/favicon.ico"}}}</script></head><body id=top><header class=header><nav class=header-nav><div class=logo><a href=https://pub.sqrt.fr/vincent/publish-assistant/ accesskey=h title="Publish Assistant (Alt + H)">Publish Assistant</a>
|
||
<span class=logo-sep>/</span>
|
||
<a class=logo-topic href=/vincent/publish-assistant/cloud-edge/ title="Edge and Cloud Systems">Edge and Cloud Systems</a><div class=logo-switches><button id=theme-toggle class=theme-toggle accesskey=t title="(Alt + T)" aria-label="Toggle theme">
|
||
<svg class="moon" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1111.21 3 7 7 0 0021 12.79z"/></svg>
|
||
<svg class="sun" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg></button></div></div><ul id=menu class=menu><li><a href=/vincent/publish-assistant/cloud-edge/venues/ title=Venues><span>Venues</span></a></li><li><a href=/vincent/publish-assistant/cloud-edge/calendar/ title=Calendar><span>Calendar</span></a></li><li><a href=/vincent/publish-assistant/cloud-edge/digests/ title=Digests><span class=active>Digests</span></a></li></ul></nav></header><main class=main><article class=post-single><header class=post-header><nav class=breadcrumbs role=navigation aria-label=Breadcrumb><a href=/vincent/publish-assistant/cloud-edge/digests/>Digests</a>
|
||
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevron-right"><polyline points="9 18 15 12 9 6"/></svg></nav><h1 class="post-title entry-hint-parent">TPDS 2024 Digest</h1><div class=post-meta><span title='2024-01-01 00:00:00 +0000 UTC'>January 1, 2024</span> · <span>Publish Assistant</span></div></header><div class="post-content md-content"><p>12 papers selected.</p><hr><h3 id=runtime-performance-anomaly-diagnosis-in-production-hpc-systems-using-active-learning>Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning<a hidden class=anchor aria-hidden=true href=#runtime-performance-anomaly-diagnosis-in-production-hpc-systems-using-active-learning>#</a></h3><p><em>Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz <em>et al.</em></em></p><p><strong>TL;DR</strong> — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort.</p><p><strong>Why notable</strong> — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads.</p><hr><h3 id=autoddl-automatic-distributed-deep-learning-with-near-optimal-bandwidth-cost>AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost<a hidden class=anchor aria-hidden=true href=#autoddl-automatic-distributed-deep-learning-with-near-optimal-bandwidth-cost>#</a></h3><p><em>Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan <em>et al.</em></em></p><p><strong>TL;DR</strong> — AutoDDL automatically searches for the distributed DNN training strategy that minimizes communication bandwidth cost while meeting performance targets.</p><p><strong>Why notable</strong> — Combining Hoefler’s communication-model expertise with automatic strategy search, this paper is essential reading for practitioners scaling DNN training across large clusters.</p><hr><h3 id=peakfs-an-ultra-high-performance-parallel-file-system-via-computing-network-storage-co-optimization-for-hpc-applications>PeakFS: An Ultra-High Performance Parallel File System via Computing-Network-Storage Co-Optimization for HPC Applications<a hidden class=anchor aria-hidden=true href=#peakfs-an-ultra-high-performance-parallel-file-system-via-computing-network-storage-co-optimization-for-hpc-applications>#</a></h3><p><em>Yixiao Chen, Haomai Yang, Kai Lu 0002, Wenlve Huang <em>et al.</em></em></p><p><strong>TL;DR</strong> — PeakFS co-optimizes compute, network, and storage layers of a parallel file system to deliver ultra-high I/O throughput for HPC workloads.</p><p><strong>Why notable</strong> — Its holistic co-design perspective sets a new performance baseline for HPC storage and provides actionable insights for next-generation parallel file system architects.</p><hr><h3 id=formal-definitions-and-performance-comparison-of-consistency-models-for-parallel-file-systems>Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems<a hidden class=anchor aria-hidden=true href=#formal-definitions-and-performance-comparison-of-consistency-models-for-parallel-file-systems>#</a></h3><p><em>Chen Wang 0004, Kathryn M. Mohror, Marc Snir</em></p><p><strong>TL;DR</strong> — This paper formalizes consistency models used by parallel file systems and provides the first systematic empirical comparison of their performance trade-offs.</p><p><strong>Why notable</strong> — Rigorous formal treatment from Snir and Mohror clarifies long-standing ambiguities in HPC storage semantics, making it an important reference for storage system designers.</p><hr><h3 id=malleability-in-modern-hpc-systems-current-experiences-challenges-and-future-opportunities>Malleability in Modern HPC Systems: Current Experiences, Challenges, and Future Opportunities<a hidden class=anchor aria-hidden=true href=#malleability-in-modern-hpc-systems-current-experiences-challenges-and-future-opportunities>#</a></h3><p><em>Ahmad Tarraf, Martin Schreiber 0001, Alberto Cascajo, Jean-Baptiste Besnard <em>et al.</em></em></p><p><strong>TL;DR</strong> — A comprehensive survey of dynamic resource malleability in HPC, covering runtime systems, job schedulers, and application-level support with lessons from production systems.</p><p><strong>Why notable</strong> — As energy-aware and burst-resilient HPC scheduling becomes critical, this broad community-driven survey is the definitive starting point for research on malleable HPC runtimes.</p><hr><h3 id=pyxis-scheduling-mixed-tasks-in-disaggregated-datacenters>Pyxis: Scheduling Mixed Tasks in Disaggregated Datacenters<a hidden class=anchor aria-hidden=true href=#pyxis-scheduling-mixed-tasks-in-disaggregated-datacenters>#</a></h3><p><em>Sheng Qi, Chao Jin, Mosharaf Chowdhury, Zhenming Liu <em>et al.</em></em></p><p><strong>TL;DR</strong> — Pyxis is a scheduler for disaggregated datacenters that jointly manages latency-sensitive and batch tasks by exploiting flexible resource pooling across the disaggregated fabric.</p><p><strong>Why notable</strong> — It tackles one of the central open problems in cloud scheduling—multi-tenancy under disaggregation—with rigorous analysis and demonstrated gains on real workloads.</p><hr><h3 id=swift-expedited-failure-recovery-for-large-scale-dnn-training>Swift: Expedited Failure Recovery for Large-Scale DNN Training<a hidden class=anchor aria-hidden=true href=#swift-expedited-failure-recovery-for-large-scale-dnn-training>#</a></h3><p><em>Yuchen Zhong, Guangming Sheng, Juncheng Liu, Jinhui Yuan <em>et al.</em></em></p><p><strong>TL;DR</strong> — Swift dramatically reduces checkpoint and recovery overhead for large-scale DNN training by combining lightweight in-memory snapshots with selective recomputation.</p><p><strong>Why notable</strong> — As training runs on hundreds of GPUs grow longer and failures become inevitable, Swift’s fault-tolerance approach directly addresses a practical bottleneck in modern deep-learning infrastructure.</p><hr><h3 id=fastload-speeding-up-data-loading-of-both-sparse-matrix-and-vector-for-spmv-on-gpus>FastLoad: Speeding Up Data Loading of Both Sparse Matrix and Vector for SpMV on GPUs<a hidden class=anchor aria-hidden=true href=#fastload-speeding-up-data-loading-of-both-sparse-matrix-and-vector-for-spmv-on-gpus>#</a></h3><p><em>Jinyu Hu, Huizhang Luo, Hong Jiang 0001, Guoqing Xiao 0001 <em>et al.</em></em></p><p><strong>TL;DR</strong> — FastLoad optimizes the memory-access pattern for loading both the sparse matrix and the dense vector in SpMV on GPUs, yielding significant throughput improvements.</p><p><strong>Why notable</strong> — SpMV is a foundational kernel for scientific computing and graph analytics; this paper’s memory-access analysis and optimizations benefit a wide class of GPU applications.</p><hr><h3 id=klnk-expanding-page-boundaries-in-a-distributed-shared-memory-system>KLNK: Expanding Page Boundaries in a Distributed Shared Memory System<a hidden class=anchor aria-hidden=true href=#klnk-expanding-page-boundaries-in-a-distributed-shared-memory-system>#</a></h3><p><em>Yiwei Ci, Michael R. Lyu, Zhan Zhang 0002, De-Cheng Zuo <em>et al.</em></em></p><p><strong>TL;DR</strong> — KLNK extends distributed shared memory page granularity to reduce false sharing and improve throughput for irregular access patterns.</p><p><strong>Why notable</strong> — It addresses a classic but unsolved bottleneck in DSM systems with a practical, page-table-level mechanism applicable to emerging disaggregated memory architectures.</p><hr><h3 id=enabling-efficient-erasure-coding-in-disaggregated-memory-systems>Enabling Efficient Erasure Coding in Disaggregated Memory Systems<a hidden class=anchor aria-hidden=true href=#enabling-efficient-erasure-coding-in-disaggregated-memory-systems>#</a></h3><p><em>Qiliang Li, Liangliang Xu, Yongkun Li 0001, Min Lyu <em>et al.</em></em></p><p><strong>TL;DR</strong> — This paper designs an erasure-coding scheme tailored to disaggregated memory, exploiting its unique bandwidth topology to achieve fault tolerance with low overhead.</p><p><strong>Why notable</strong> — Fault tolerance in disaggregated memory is an open problem of growing importance; this work provides concrete mechanisms and strong performance results.</p><hr><h3 id=simple-fast-and-widely-applicable-concurrent-memory-reclamation-via-neutralization>Simple, Fast and Widely Applicable Concurrent Memory Reclamation via Neutralization<a hidden class=anchor aria-hidden=true href=#simple-fast-and-widely-applicable-concurrent-memory-reclamation-via-neutralization>#</a></h3><p><em>Ajay Singh 0002, Trevor Alexander Brown, Ali José Mashtizadeh</em></p><p><strong>TL;DR</strong> — Neutralization is a new mechanism for safe memory reclamation in lock-free data structures that is simpler, faster, and more portable than prior approaches.</p><p><strong>Why notable</strong> — Safe memory reclamation is a pervasive challenge in concurrent programming; this algorithm’s breadth of applicability and performance improvements make it highly reusable.</p><hr><h3 id=deeptm-efficient-tensor-management-in-heterogeneous-memory-for-dnn-training>DeepTM: Efficient Tensor Management in Heterogeneous Memory for DNN Training<a hidden class=anchor aria-hidden=true href=#deeptm-efficient-tensor-management-in-heterogeneous-memory-for-dnn-training>#</a></h3><p><em>Haoran Zhou, Wei Rang, Hongyang Chen 0001, Xiaobo Zhou 0002 <em>et al.</em></em></p><p><strong>TL;DR</strong> — DeepTM dynamically manages tensor placement across DRAM and NVM during DNN training to reduce memory pressure and improve throughput.</p><p><strong>Why notable</strong> — As model sizes outpace GPU memory, heterogeneous memory management becomes critical; DeepTM provides a practical, training-aware solution with measurable benefits.</p></div><footer class=post-footer><ul class=post-tags></ul><nav class=paginav><a class=prev href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/tocs-2024/><span class=title>« Prev</span>
|
||
<span>TOCS 2024 Digest</span></a></nav></footer></article></main><footer class=footer><span>© 2026 <a href=https://pub.sqrt.fr/vincent/publish-assistant/>Publish Assistant</a></span> ·
|
||
<span>Powered by
|
||
<a href="https://gohugo.io/?utm_source=papermod" rel=noopener target=_blank>Hugo</a> &
|
||
<a href=https://github.com/adityatelange/hugo-PaperMod/ rel=noopener target=_blank>PaperMod</a></span></footer><a href=#top id=top-link class="top-link hidden" aria-label="go to top" title="Go to Top (Alt + G)" accesskey=g><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevrons-up"><polyline points="17 11 12 6 7 11"/><polyline points="17 18 12 13 7 18"/></svg>
|
||
</a><script>let menu=document.getElementById("menu");if(menu){const e=localStorage.getItem("menu-scroll-position");e&&(menu.scrollLeft=parseInt(e,10)),menu.onscroll=function(){localStorage.setItem("menu-scroll-position",menu.scrollLeft)}}document.querySelectorAll('a[href^="#"]').forEach(e=>{e.addEventListener("click",function(e){e.preventDefault();var t=this.getAttribute("href").substr(1);window.matchMedia("(prefers-reduced-motion: reduce)").matches?document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView():document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView({behavior:"smooth"}),t==="top"?history.replaceState(null,null," "):history.pushState(null,null,`#${t}`)})})</script><script>var toplink=document.getElementById("top-link");window.onscroll=function(){const e=window.innerHeight;document.body.scrollTop>e||document.documentElement.scrollTop>e?toplink.classList.remove("hidden"):toplink.classList.add("hidden")}</script><script>document.getElementById("theme-toggle").addEventListener("click",()=>{const e=document.querySelector("html");e.dataset.theme==="dark"?(e.dataset.theme="light",localStorage.setItem("pref-theme","light")):(e.dataset.theme="dark",localStorage.setItem("pref-theme","dark"))})</script></body></html> |