tl;drs for the top papers.

HPDC 2025 Digest

10 papers selected. Parameterized Algorithms for Non-uniform All-to-all Ke Fan, Jens Domke, Seydou Ba, Sidharth Kumar TL;DR — Introduces parameterized algorithms that adapt non-uniform all-to-all collective communication to heterogeneous network topologies, reducing message contention and improving throughput. Why notable — Non-uniform all-to-all is a performance bottleneck in many HPC applications; topology-aware parameterization directly benefits MPI implementations on dragonfly and fat-tree networks at scale. → Read paper DPU-KV: On the Benefits of DPU Offloading for In-Memory Key-Value Stores at the Edge Arjun Kashyap, Yuke Li 0003, Xiaoyi Lu 0001 ...

July 20, 2025 · Publish Assistant

ATC 2025 Digest

13 papers selected. ASTERINAS: A Linux ABI-Compatible, Rust-Based Framekernel OS with a Small and Sound TCB Yuke Peng, Hongliang Tian, Junyang Zhang, Ruihan Li et al. TL;DR — A production-grade OS kernel written in Rust that exposes a full Linux ABI while confining unsafe code to a small, formally-audited framekernel core. Why notable — ASTERINAS demonstrates that Linux compatibility and memory-safety guarantees are not mutually exclusive — unsafe Rust is isolated to under 5 kloc of framework code, giving systems operators a credible path toward a safer Linux-compatible kernel without sacrificing application portability. ...

July 9, 2025 · Publish Assistant

MobiSys 2025 Digest

12 papers selected. Hopter: a Safe, Robust, and Responsive Embedded Operating System Zhiyao Ma, Guojun Chen, Zhuo Chen 0011, Lin Zhong 0001 TL;DR — Hopter is a new embedded OS that enforces memory safety and real-time responsiveness through a Rust-based task model with cooperative and preemptive scheduling co-designed from the ground up. Why notable — Building a ground-up safe embedded OS is a long-standing challenge; Hopter addresses it without sacrificing the determinism that IoT and robotics workloads demand, offering a credible alternative to unsafe C-based RTOSes. ...

June 23, 2025 · Publish Assistant

IPDPS 2025 Digest

12 papers selected. Enhancing OmpSs-2 Suspendable Tasks by Combining Operating System and User-Level Threads with C++ Coroutines Arnau Cinca, Aleix Roca, Kevin Sala, Raúl Peñacoba Veigas et al. TL;DR — Extends the OmpSs-2 task-based runtime with C++ coroutines to implement suspendable tasks that can yield while blocked on I/O or communication without stalling the OS thread. Why notable — Suspendable tasks are a key missing primitive for overlapping computation and communication in task-graph runtimes; the hybrid OS/user-level thread design avoids the overhead of full context switches while remaining portable, with broad implications for OpenMP-style programming on modern heterogeneous nodes. ...

May 19, 2025 · Publish Assistant

NSDI 2025 Digest

13 papers selected. PRED: Performance-oriented Random Early Detection for Consistently Stable Performance in Datacenters Xinle Du, Tong Li 0014, Guangmeng Zhou, Zhuotao Liu et al. TL;DR — PRED redesigns AQM by making drop probability a direct function of per-flow performance targets rather than queue length, eliminating the instability of classic RED in modern datacenter workloads. Why notable — RED has been a cornerstone of congestion control for decades; PRED’s performance-centric reformulation challenges a long-held design axiom and demonstrates significantly lower tail latency at scale. It opens the door to intent-driven AQM as a first-class primitive in datacenter switches. ...

April 28, 2025 · Publish Assistant

EuroSys 2025 Digest

13 papers selected. Empowering WebAssembly with Thin Kernel Interfaces Arjun Ramesh, Tianshu Huang, Ben L. Titzer, Anthony Rowe 0001 TL;DR — A new OS interface design exposes thin, capability-based kernel primitives directly to WebAssembly modules, eliminating the POSIX translation layer. Why notable — WebAssembly is increasingly used beyond the browser as a portable, sandboxed compute substrate; this work shows that rethinking the system interface from scratch yields significantly lower overhead and better safety properties than layering Wasm on top of POSIX. ...

March 30, 2025 · Publish Assistant

FGCS 2025 Digest

12 papers selected. Multifacets of lossy compression for scientific data in the Joint-Laboratory of Extreme Scale Computing Franck Cappello, Mario C. Acosta, Emmanuel Agullo, Hartwig Anzt et al. TL;DR — A joint JLESC survey covering error-bounded lossy compressors (SZ, ZFP, MGARD) across simulation, AI, and in-situ analytics use cases, with benchmarks on real scientific datasets at extreme scale. Why notable — The most comprehensive cross-site evaluation of scientific data compression to date, providing actionable guidance on compressor selection for different numerical kernels and accuracy requirements. ...

January 1, 2025 · Publish Assistant

IC 2025 Digest

10 papers selected. Rethinking Computing Systems in the Era of Climate Crisis: A Call for a Sustainable Computing Continuum Ella Peltonen, Suzan Bayhan, David Bermbach, Sebastian Buschjäger et al. TL;DR — A multi-author position paper calling for carbon-aware design principles across the cloud-to-edge computing continuum, surveying energy measurement, workload scheduling, and hardware lifecycle challenges. Why notable — Establishes a community research agenda for sustainable computing infrastructure at a time when datacenter and edge energy consumption is under increasing regulatory and societal scrutiny. ...

January 1, 2025 · Publish Assistant

JPDC 2025 Digest

12 papers selected. Throughput of Byzantine Broadcast Ruomu Hou, Haifeng Yu, Prateek Saxena TL;DR — Establishes tight throughput bounds for Byzantine broadcast protocols and constructs algorithms that saturate those bounds, separating throughput from latency in the fault-tolerant broadcast landscape. Why notable — Provides the first rigorous throughput characterization of Byzantine broadcast, a fundamental primitive whose capacity limits were previously unquantified. How to reduce the number of steps for (multi-valued validated) Byzantine agreement? Baohan Huang, Haibin Zhang, Chao Liu 0039, Shengli Liu 0001 et al. ...

January 1, 2025 · Publish Assistant

Middleware 2025 Digest

12 papers selected. Recipe: Hardware-Accelerated Replication Protocols: Rethinking Crash Fault Tolerance Protocols for Untrusted Cloud Environments Dimitra Giantsidi, Emmanouil Giortamis, Julian Pritzi, Maurice Bailleu et al. TL;DR — Redesigns crash fault-tolerance protocols using hardware acceleration (TEEs/SmartNICs) to deliver replication with strong guarantees in untrusted cloud environments. Efficient Performance Guarantees for Function-as-a-Service with Cloud Allocators Hai Duc Nguyen 0005, Andrew A. Chien TL;DR — Introduces cloud-allocator abstractions that provide formal performance guarantees for serverless functions, addressing the unpredictability of shared FaaS infrastructure. ...

January 1, 2025 · Publish Assistant

OSDI 2025 Digest

13 papers selected. Basilisk: Using Provenance Invariants to Automate Proofs of Undecidable Protocols Tony Nuda Zhang, Keshav Singh, Tej Chajed, Manos Kapritsos et al. TL;DR — Automates the construction of correctness proofs for distributed protocols that were previously considered undecidable, advancing the state of the art in verified systems. Mako: Speculative Distributed Transactions with Geo-Replication Weihai Shen, Yang Cui, Siddhartha Sen 0001, Sebastian Angel et al. TL;DR — Combines speculative execution with geo-replication to deliver low-latency distributed transactions without sacrificing consistency, addressing a fundamental tension in wide-area systems. ...

January 1, 2025 · Publish Assistant

SC 2025 Digest

15 papers selected. Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban M. Rangel et al. TL;DR — Delivers the first trillion-particle cosmological hydrodynamics simulation on exascale hardware, demonstrating sustained petaflop-scale performance on a flagship scientific application. Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka et al. TL;DR — Achieves sustained exascale performance for first-principles quantum transport at 42,240 atoms, establishing a new scale record for the GW many-body perturbation method. ...

January 1, 2025 · Publish Assistant

SEC 2025 Digest

12 papers selected. lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models Haoxin Wang 0003 TL;DR — Provides the first detailed runtime profiling framework for on-device LLM inference, revealing key latency bottlenecks across diverse edge hardware configurations. SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving Xiangchen Li TL;DR — Adapts speculative decoding to edge serving constraints, reducing LLM token generation latency while respecting the tight memory and compute budgets of edge nodes. ...

January 1, 2025 · Publish Assistant

SoCC 2025 Digest

13 papers selected. From Bottleneck to Breakthrough: Optimizing Scheduling for Hyperscale Containerized Clusters Bing Li, Yuquan Ren, Xinyi Song, Zhilei Liu et al. TL;DR — Documents production-scale scheduling improvements at a hyperscale cloud provider, demonstrating how targeted optimizations reduce scheduling tail latency and increase cluster utilization in real containerized workloads. CPU-Limits kill Performance: Time to rethink Resource Control Chirag C. Shetty, Sarthak Chakraborty, Hubertus Franke, Larisa Shwartz et al. TL;DR — Challenges the conventional use of CPU cgroup limits in cloud environments, showing through production evidence that CFS bandwidth throttling degrades application QoS and proposing a rethink of resource control abstractions. ...

January 1, 2025 · Publish Assistant

SOSP 2025 Digest

14 papers selected. LithOS: An Operating System for Efficient Machine Learning on GPUs Patrick H. Coppock, Brian Zhang, Eliot H. Solomon, Vasilis Kypriotis et al. TL;DR — Designs a dedicated OS for GPU ML workloads, rethinking scheduling and resource management at the kernel level for accelerator-centric computing. CHERIoT RTOS: An OS for Fine-Grained Memory-Safe Compartments on Low-Cost Embedded Devices Saar Amar, Tony Chen, David Chisnall, Nathaniel Wesley Filardo et al. ...

January 1, 2025 · Publish Assistant

TC 2025 Digest

12 papers selected. RV-CURE: A RISC-V Capability Architecture for Full Memory Safety Yonghae Kim, Anurag Kar, Jaewon Lee, Jaekyu Lee et al. TL;DR — Extends the RISC-V ISA with hardware capabilities to enforce full memory safety—including bounds checking and pointer provenance—across the entire software stack. Why notable — Demonstrates that capability-based memory safety can be integrated into an open ISA at low cost, with implications for deploying safe-by-default embedded and server systems. ...

January 1, 2025 · Publish Assistant

TCC 2025 Digest

12 papers selected. DRKC: Deep Reinforcement Learning Enhanced Microservice Scheduling on Kubernetes Clusters in Cloud-Edge Environment Jian Jiang, Qianmu Li, Pengchuan Wang, Yunhuai Liu TL;DR — DRKC uses deep reinforcement learning to schedule microservices across Kubernetes clusters spanning cloud and edge nodes, optimizing latency and resource utilization. Why notable — One of the few papers to tackle DRL-based microservice placement at the Kubernetes level in a real cloud-edge topology, making it directly actionable for practitioners. ...

January 1, 2025 · Publish Assistant

TOCS 2025 Digest

10 papers selected. Whole-system Persistence Made Efficient with Tree-structured Checkpointing on Microkernel Mingkai Dong, Fangnuo Wu, Gequan Mo, Haibo Chen TL;DR — A microkernel-based whole-system persistence scheme uses tree-structured incremental checkpointing to achieve low-overhead, crash-consistent snapshots of the entire OS state. Why notable — Whole-system persistence is a foundational building block for reliable systems; this paper shows it can be done efficiently within a microkernel architecture. XpuTEE: A High-Performance and Practical Heterogeneous Trusted Execution Environment for GPUs Shulin Fan, Zhichao Hua, Yubin Xia, Haibo Chen ...

January 1, 2025 · Publish Assistant

TPDS 2025 Digest

12 papers selected. HARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management Kyrian Adimora, Hongyang Sun 0001 TL;DR — HARMONIC applies uncertainty-aware multi-objective optimization to jointly minimize energy consumption and maximize performance for HPC resource allocation. Why notable — It directly addresses the growing demand for energy-proportional HPC scheduling with principled probabilistic models, making it relevant to both system designers and green-computing researchers. MIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems Yiqin Dai, Ruibo Wang, Yong Dong, Min Xie et al. ...

January 1, 2025 · Publish Assistant

Middleware 2024 Digest

11 papers selected. Chasing Lightspeed Consensus: Fast Wide-Area Byzantine Replication with Mercury Christian Berger 0006, Lívio Rodrigues, Hans P. Reiser, Vinicius Vielmo Cogo et al. TL;DR — Mercury is a wide-area Byzantine fault-tolerant replication protocol that minimises latency by exploiting geographic locality and pipelining to approach the theoretical lightspeed bound. Why notable — Achieving near-lightspeed latency in Byzantine replication across wide-area networks has been a long-standing open challenge; Mercury’s design demonstrates it is practically attainable. The result raises the bar for what production BFT middleware can deliver in geo-distributed deployments. ...

December 2, 2024 · Publish Assistant