Files
publish-assistant/cloud-edge/digests/sc-2025/index.html
2026-08-18 13:39:21 +00:00

29 lines
27 KiB
HTML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!doctype html><html lang=en dir=auto data-theme=auto><head><meta charset=utf-8><meta http-equiv=X-UA-Compatible content="IE=edge"><meta name=viewport content="width=device-width,initial-scale=1,shrink-to-fit=no"><meta name=robots content="index, follow"><title>SC 2025 Digest | Publish Assistant</title><meta name=keywords content><meta name=description content="15 papers selected.
Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability
Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban M. Rangel et al.
TL;DR — Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application.
Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance
Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka et al.
TL;DR — Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method."><meta name=author content="Publish Assistant"><link rel=canonical href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/sc-2025/><link crossorigin=anonymous href=/vincent/publish-assistant/assets/css/stylesheet.d72f07832e13c592b3edba91680bfe70f01daac396179bcace0ac36e8e0494c6.css integrity="sha256-1y8Hgy4TxZKz7bqRaAv+cPAdqsOWF5vKzgrDbo4ElMY=" rel="preload stylesheet" as=style><link rel=icon href=https://pub.sqrt.fr/vincent/publish-assistant/favicon.ico><link rel=icon type=image/png sizes=16x16 href=https://pub.sqrt.fr/vincent/publish-assistant/favicon-16x16.png><link rel=icon type=image/png sizes=32x32 href=https://pub.sqrt.fr/vincent/publish-assistant/favicon-32x32.png><link rel=apple-touch-icon href=https://pub.sqrt.fr/vincent/publish-assistant/apple-touch-icon.png><link rel=mask-icon href=https://pub.sqrt.fr/vincent/publish-assistant/safari-pinned-tab.svg><meta name=theme-color content="#2e2e33"><meta name=msapplication-TileColor content="#2e2e33"><link rel=alternate hreflang=en href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/sc-2025/><noscript><style>#theme-toggle,.top-link{display:none}</style><style>@media(prefers-color-scheme:dark){:root{--theme:rgb(29, 30, 32);--entry:rgb(46, 46, 51);--primary:rgb(218, 218, 219);--secondary:rgb(155, 156, 157);--tertiary:rgb(65, 66, 68);--content:rgb(196, 196, 197);--code-block-bg:rgb(46, 46, 51);--code-bg:rgb(55, 56, 62);--border:rgb(51, 51, 51);color-scheme:dark}.list{background:var(--theme)}.toc{background:var(--entry)}}</style></noscript><script>localStorage.getItem("pref-theme")==="dark"?document.querySelector("html").dataset.theme="dark":localStorage.getItem("pref-theme")==="light"?document.querySelector("html").dataset.theme="light":window.matchMedia("(prefers-color-scheme: dark)").matches?document.querySelector("html").dataset.theme="dark":document.querySelector("html").dataset.theme="light"</script><meta property="og:url" content="https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/sc-2025/"><meta property="og:site_name" content="Publish Assistant"><meta property="og:title" content="SC 2025 Digest"><meta property="og:description" content="15 papers selected.
Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban M. Rangel et al.
TL;DR — Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application.
Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka et al.
TL;DR — Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method."><meta property="og:locale" content="en_us"><meta property="og:type" content="article"><meta property="article:section" content="cloud-edge"><meta property="article:published_time" content="2025-01-01T00:00:00+00:00"><meta property="article:modified_time" content="2025-01-01T00:00:00+00:00"><meta name=twitter:card content="summary"><meta name=twitter:title content="SC 2025 Digest"><meta name=twitter:description content="15 papers selected.
Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban M. Rangel et al.
TL;DR — Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application.
Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka et al.
TL;DR — Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method."><script type=application/ld+json>{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Edge and Cloud Systems","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/"},{"@type":"ListItem","position":2,"name":"Digests","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/"},{"@type":"ListItem","position":3,"name":"SC 2025 Digest","item":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/sc-2025/"}]}</script><script type=application/ld+json>{"@context":"https://schema.org","@type":"BlogPosting","headline":"SC 2025 Digest","name":"SC 2025 Digest","description":"15 papers selected.\nCosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban M. Rangel et al.\nTL;DR — Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application.\nAb-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka et al.\nTL;DR — Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method.\n","keywords":[],"articleBody":"15 papers selected.\nCosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban M. Rangel et al.\nTL;DR — Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application.\nAb-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka et al.\nTL;DR — Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method.\nKilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers Kai Xu, Maoxue Yu, Yuhu Chen, Jie Gao et al.\nTL;DR — Integrates AI acceleration into a kilometer-scale climate model to reach year-scale simulation throughput, showing how MLphysics hybrid approaches can redefine climate modeling at supercomputer scale.\nUno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini et al.\nTL;DR — Proposes a unified congestion control and reliable transport architecture spanning intra- and inter-datacenter links, with strong throughput and latency results relevant to AI and HPC clusters.\nSDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication Mikhail Khalilov, Siyuan Shen, Marcin Chrapek, Tiancheng Chen et al.\nTL;DR — Introduces software-defined reliability for RDMA at global scale, decoupling reliability policies from hardware to dramatically improve fault tolerance and reconfigurability in large-scale HPC networks.\nBine Trees: Enhancing Collective Operations by Optimizing Communication Locality Daniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli, Tommaso Bonato et al.\nTL;DR — Presents bine tree topologies for MPI collective operations that exploit communication locality, yielding significant latency and bandwidth improvements over standard binomial trees on modern HPC interconnects.\nSTELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems Chris Egersdoerfer, Philip H. Carns, Shane Snyder, Robert Ross et al.\nTL;DR — Demonstrates that an LLM-driven autonomous reasoning engine can tune parallel file system parameters as effectively as expert hand-tuning, opening a new direction for self-optimizing HPC storage.\nPhoenix: A Refactored I/O Stack for GPU Direct Storage without Phony Buffers Jianqin Yan, Shi Qiu 0012, Yina Lv, Yifan Hu et al.\nTL;DR — Redesigns the GPU direct storage I/O stack to eliminate staging buffers, achieving large bandwidth gains for GPU-to-SSD transfers critical to LLM training and scientific data workflows.\nBreaking the System Noise Barrier at Exascale Edgar A. León, Joseph Glenski, Mark J. Stock, Kim H. McMahon et al.\nTL;DR — Provides a rigorous characterization and mitigation of OS and hardware noise at exascale, demonstrating measurable improvements in collective communication performance on a real production system.\nStory of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs Shengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan et al.\nTL;DR — Delivers the first detailed side-by-side hardware fault-injection study of H100 and A100 GPUs, revealing how architecture changes in Hopper alter error propagation and resilience for HPC and AI workloads.\nExploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems Pengfei Yu 0002, Jingjing Gu, Hao Han, Dazhong Shen et al.\nTL;DR — Characterizes real-world failure modes of large-scale LLM training on HPC clusters and proposes targeted mitigation strategies, providing essential reliability insights for AI infrastructure operators.\nXaaS Containers: Performance-Portable Representation With Source and IR Containers Marcin Copik, Eiman Alnuaimi, Alok Kamatar, Valérie Hayot-Sasson et al.\nTL;DR — Proposes source- and IR-level HPC containers that enable performance portability across heterogeneous architectures without recompilation, addressing a key deployment challenge for reproducible HPC software.\ncMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications Xi Wang 0027, Bin Ma, Jongryool Kim, Byungil Koh et al.\nTL;DR — Exploits CXL memory semantics to implement MPI communication primitives with dramatically reduced software overhead, demonstrating a promising path for memory-centric supercomputer interconnects.\nX-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms Yueming Yuan, Ahan Gupta, Jianping Li, Sajal Dash et al.\nTL;DR — Addresses the communication and load-balance bottlenecks of sparse Mixture-of-Experts training at scale, achieving efficient utilization of large GPU clusters for next-generation LLM workloads.\nBenchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing Oscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding 0006 et al.\nTL;DR — Develops fine-grained benchmark-driven energy models for GPU supercomputers that attribute power consumption to individual components and workloads, enabling principled energy optimization at the facility level.\n","wordCount":"734","inLanguage":"en","datePublished":"2025-01-01T00:00:00Z","dateModified":"2025-01-01T00:00:00Z","author":{"@type":"Person","name":"Publish Assistant"},"mainEntityOfPage":{"@type":"WebPage","@id":"https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/sc-2025/"},"publisher":{"@type":"Organization","name":"Publish Assistant","logo":{"@type":"ImageObject","url":"https://pub.sqrt.fr/vincent/publish-assistant/favicon.ico"}}}</script></head><body id=top><header class=header><nav class=header-nav><div class=logo><a href=https://pub.sqrt.fr/vincent/publish-assistant/ accesskey=h title="Publish Assistant (Alt + H)">Publish Assistant</a>
<span class=logo-sep>/</span>
<a class=logo-topic href=/vincent/publish-assistant/cloud-edge/ title="Edge and Cloud Systems">Edge and Cloud Systems</a><div class=logo-switches><button id=theme-toggle class=theme-toggle accesskey=t title="(Alt + T)" aria-label="Toggle theme">
<svg class="moon" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1111.21 3 7 7 0 0021 12.79z"/></svg>
<svg class="sun" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg></button></div></div><ul id=menu class=menu><li><a href=/vincent/publish-assistant/cloud-edge/venues/ title=Venues><span>Venues</span></a></li><li><a href=/vincent/publish-assistant/cloud-edge/calendar/ title=Calendar><span>Calendar</span></a></li><li><a href=/vincent/publish-assistant/cloud-edge/digests/ title=Digests><span class=active>Digests</span></a></li></ul></nav></header><main class=main><article class=post-single><header class=post-header><nav class=breadcrumbs role=navigation aria-label=Breadcrumb><a href=/vincent/publish-assistant/cloud-edge/digests/>Digests</a>
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevron-right"><polyline points="9 18 15 12 9 6"/></svg></nav><h1 class="post-title entry-hint-parent">SC 2025 Digest</h1><div class=post-meta><span title='2025-01-01 00:00:00 +0000 UTC'>January 1, 2025</span>&nbsp;·&nbsp;<span>Publish Assistant</span></div></header><div class="post-content md-content"><p>15 papers selected.</p><hr><h3 id=cosmological-hydrodynamics-at-exascale-a-trillion-particle-leap-in-capability>Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability<a hidden class=anchor aria-hidden=true href=#cosmological-hydrodynamics-at-exascale-a-trillion-particle-leap-in-capability>#</a></h3><p><em>Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban M. Rangel <em>et al.</em></em></p><p><strong>TL;DR</strong> — Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application.</p><hr><h3 id=ab-initio-quantum-transport-with-the-gw-approximation-42-240-atoms-and-sustained-exascale-performance>Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance<a hidden class=anchor aria-hidden=true href=#ab-initio-quantum-transport-with-the-gw-approximation-42-240-atoms-and-sustained-exascale-performance>#</a></h3><p><em>Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka <em>et al.</em></em></p><p><strong>TL;DR</strong> — Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method.</p><hr><h3 id=kilometer-scale-ai-powered-and-performance-portable-earth-system-model-ap3esm-to-achieve-year-scale-simulation-speed-on-heterogeneous-supercomputers>Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers<a hidden class=anchor aria-hidden=true href=#kilometer-scale-ai-powered-and-performance-portable-earth-system-model-ap3esm-to-achieve-year-scale-simulation-speed-on-heterogeneous-supercomputers>#</a></h3><p><em>Kai Xu, Maoxue Yu, Yuhu Chen, Jie Gao <em>et al.</em></em></p><p><strong>TL;DR</strong> — Integrates AI acceleration into a kilometer-scale climate model to reach year-scale simulation throughput, showing how MLphysics hybrid approaches can redefine climate modeling at supercomputer scale.</p><hr><h3 id=uno-a-one-stop-solution-for-inter--and-intra-data-center-congestion-control-and-reliable-connectivity>Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity<a hidden class=anchor aria-hidden=true href=#uno-a-one-stop-solution-for-inter--and-intra-data-center-congestion-control-and-reliable-connectivity>#</a></h3><p><em>Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini <em>et al.</em></em></p><p><strong>TL;DR</strong> — Proposes a unified congestion control and reliable transport architecture spanning intra- and inter-datacenter links, with strong throughput and latency results relevant to AI and HPC clusters.</p><hr><h3 id=sdr-rdma-software-defined-reliability-architecture-for-planetary-scale-rdma-communication>SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication<a hidden class=anchor aria-hidden=true href=#sdr-rdma-software-defined-reliability-architecture-for-planetary-scale-rdma-communication>#</a></h3><p><em>Mikhail Khalilov, Siyuan Shen, Marcin Chrapek, Tiancheng Chen <em>et al.</em></em></p><p><strong>TL;DR</strong> — Introduces software-defined reliability for RDMA at global scale, decoupling reliability policies from hardware to dramatically improve fault tolerance and reconfigurability in large-scale HPC networks.</p><hr><h3 id=bine-trees-enhancing-collective-operations-by-optimizing-communication-locality>Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality<a hidden class=anchor aria-hidden=true href=#bine-trees-enhancing-collective-operations-by-optimizing-communication-locality>#</a></h3><p><em>Daniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli, Tommaso Bonato <em>et al.</em></em></p><p><strong>TL;DR</strong> — Presents bine tree topologies for MPI collective operations that exploit communication locality, yielding significant latency and bandwidth improvements over standard binomial trees on modern HPC interconnects.</p><hr><h3 id=stellar-storage-tuning-engine-leveraging-llm-autonomous-reasoning-for-high-performance-parallel-file-systems>STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems<a hidden class=anchor aria-hidden=true href=#stellar-storage-tuning-engine-leveraging-llm-autonomous-reasoning-for-high-performance-parallel-file-systems>#</a></h3><p><em>Chris Egersdoerfer, Philip H. Carns, Shane Snyder, Robert Ross <em>et al.</em></em></p><p><strong>TL;DR</strong> — Demonstrates that an LLM-driven autonomous reasoning engine can tune parallel file system parameters as effectively as expert hand-tuning, opening a new direction for self-optimizing HPC storage.</p><hr><h3 id=phoenix-a-refactored-io-stack-for-gpu-direct-storage-without-phony-buffers>Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony Buffers<a hidden class=anchor aria-hidden=true href=#phoenix-a-refactored-io-stack-for-gpu-direct-storage-without-phony-buffers>#</a></h3><p><em>Jianqin Yan, Shi Qiu 0012, Yina Lv, Yifan Hu <em>et al.</em></em></p><p><strong>TL;DR</strong> — Redesigns the GPU direct storage I/O stack to eliminate staging buffers, achieving large bandwidth gains for GPU-to-SSD transfers critical to LLM training and scientific data workflows.</p><hr><h3 id=breaking-the-system-noise-barrier-at-exascale>Breaking the System Noise Barrier at Exascale<a hidden class=anchor aria-hidden=true href=#breaking-the-system-noise-barrier-at-exascale>#</a></h3><p><em>Edgar A. León, Joseph Glenski, Mark J. Stock, Kim H. McMahon <em>et al.</em></em></p><p><strong>TL;DR</strong> — Provides a rigorous characterization and mitigation of OS and hardware noise at exascale, demonstrating measurable improvements in collective communication performance on a real production system.</p><hr><h3 id=story-of-two-gpus-characterizing-the-resilience-of-hopper-h100-and-ampere-a100-gpus>Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs<a hidden class=anchor aria-hidden=true href=#story-of-two-gpus-characterizing-the-resilience-of-hopper-h100-and-ampere-a100-gpus>#</a></h3><p><em>Shengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan <em>et al.</em></em></p><p><strong>TL;DR</strong> — Delivers the first detailed side-by-side hardware fault-injection study of H100 and A100 GPUs, revealing how architecture changes in Hopper alter error propagation and resilience for HPC and AI workloads.</p><hr><h3 id=exploring-and-mitigating-failure-behavior-of-large-language-model-training-workloads-in-hpc-systems>Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems<a hidden class=anchor aria-hidden=true href=#exploring-and-mitigating-failure-behavior-of-large-language-model-training-workloads-in-hpc-systems>#</a></h3><p><em>Pengfei Yu 0002, Jingjing Gu, Hao Han, Dazhong Shen <em>et al.</em></em></p><p><strong>TL;DR</strong> — Characterizes real-world failure modes of large-scale LLM training on HPC clusters and proposes targeted mitigation strategies, providing essential reliability insights for AI infrastructure operators.</p><hr><h3 id=xaas-containers-performance-portable-representation-with-source-and-ir-containers>XaaS Containers: Performance-Portable Representation With Source and IR Containers<a hidden class=anchor aria-hidden=true href=#xaas-containers-performance-portable-representation-with-source-and-ir-containers>#</a></h3><p><em>Marcin Copik, Eiman Alnuaimi, Alok Kamatar, Valérie Hayot-Sasson <em>et al.</em></em></p><p><strong>TL;DR</strong> — Proposes source- and IR-level HPC containers that enable performance portability across heterogeneous architectures without recompilation, addressing a key deployment challenge for reproducible HPC software.</p><hr><h3 id=cmpi-using-cxl-memory-sharing-for-mpi-one-sided-and-two-sided-inter-node-communications>cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications<a hidden class=anchor aria-hidden=true href=#cmpi-using-cxl-memory-sharing-for-mpi-one-sided-and-two-sided-inter-node-communications>#</a></h3><p><em>Xi Wang 0027, Bin Ma, Jongryool Kim, Byungil Koh <em>et al.</em></em></p><p><strong>TL;DR</strong> — Exploits CXL memory semantics to implement MPI communication primitives with dramatically reduced software overhead, demonstrating a promising path for memory-centric supercomputer interconnects.</p><hr><h3 id=x-moe-enabling-scalable-training-for-emerging-mixture-of-experts-architectures-on-hpc-platforms>X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms<a hidden class=anchor aria-hidden=true href=#x-moe-enabling-scalable-training-for-emerging-mixture-of-experts-architectures-on-hpc-platforms>#</a></h3><p><em>Yueming Yuan, Ahan Gupta, Jianping Li, Sajal Dash <em>et al.</em></em></p><p><strong>TL;DR</strong> — Addresses the communication and load-balance bottlenecks of sparse Mixture-of-Experts training at scale, achieving efficient utilization of large GPU clusters for next-generation LLM workloads.</p><hr><h3 id=benchmark-driven-models-for-energy-analysis-and-attribution-of-gpu-accelerated-supercomputing>Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing<a hidden class=anchor aria-hidden=true href=#benchmark-driven-models-for-energy-analysis-and-attribution-of-gpu-accelerated-supercomputing>#</a></h3><p><em>Oscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding 0006 <em>et al.</em></em></p><p><strong>TL;DR</strong> — Develops fine-grained benchmark-driven energy models for GPU supercomputers that attribute power consumption to individual components and workloads, enabling principled energy optimization at the facility level.</p></div><footer class=post-footer><ul class=post-tags></ul><nav class=paginav><a class=prev href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/osdi-2025/><span class=title>« Prev</span>
<span>OSDI 2025 Digest</span>
</a><a class=next href=https://pub.sqrt.fr/vincent/publish-assistant/cloud-edge/digests/sec-2025/><span class=title>Next »</span>
<span>SEC 2025 Digest</span></a></nav></footer></article></main><footer class=footer><span>&copy; 2026 <a href=https://pub.sqrt.fr/vincent/publish-assistant/>Publish Assistant</a></span> ·
<span>Powered by
<a href="https://gohugo.io/?utm_source=papermod" rel=noopener target=_blank>Hugo</a> &
<a href=https://github.com/adityatelange/hugo-PaperMod/ rel=noopener target=_blank>PaperMod</a></span></footer><a href=#top id=top-link class="top-link hidden" aria-label="go to top" title="Go to Top (Alt + G)" accesskey=g><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevrons-up"><polyline points="17 11 12 6 7 11"/><polyline points="17 18 12 13 7 18"/></svg>
</a><script>let menu=document.getElementById("menu");if(menu){const e=localStorage.getItem("menu-scroll-position");e&&(menu.scrollLeft=parseInt(e,10)),menu.onscroll=function(){localStorage.setItem("menu-scroll-position",menu.scrollLeft)}}document.querySelectorAll('a[href^="#"]').forEach(e=>{e.addEventListener("click",function(e){e.preventDefault();var t=this.getAttribute("href").substr(1);window.matchMedia("(prefers-reduced-motion: reduce)").matches?document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView():document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView({behavior:"smooth"}),t==="top"?history.replaceState(null,null," "):history.pushState(null,null,`#${t}`)})})</script><script>var toplink=document.getElementById("top-link");window.onscroll=function(){const e=window.innerHeight;document.body.scrollTop>e||document.documentElement.scrollTop>e?toplink.classList.remove("hidden"):toplink.classList.add("hidden")}</script><script>document.getElementById("theme-toggle").addEventListener("click",()=>{const e=document.querySelector("html");e.dataset.theme==="dark"?(e.dataset.theme="light",localStorage.setItem("pref-theme","light")):(e.dataset.theme="dark",localStorage.setItem("pref-theme","dark"))})</script></body></html>