tl;drs for the top papers.

SOSP 2024 Digest

13 papers selected. Verus: A Practical Foundation for Systems Verification Andrea Lattuada 0001, Travis Hance, Jay Bosamiya, Matthias Brun 0002 et al. TL;DR — Verus is a Rust-based verification framework that makes formal proofs of low-level systems code tractable at scale, covering memory safety, functional correctness, and concurrency. Why notable — Formal verification of real systems code has long been impractical; Verus closes the usability gap by integrating SMT-based proofs directly into a systems programming language, making it the most broadly applicable verification tool for the OS community to date. ...

November 5, 2024 · Publish Assistant

SoCC 2024 Digest

12 papers selected. Queue Management for SLO-Oriented Large Language Model Serving Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu et al. TL;DR — A queue management framework that enforces latency SLOs for LLM serving by dynamically routing and prioritizing requests across heterogeneous inference capacity. Why notable — As LLM deployments move into production clouds, meeting strict time-to-first-token and total latency SLOs becomes critical; this work directly addresses that gap with a practical, deployable solution. It is one of the first papers to treat LLM serving as a cloud SLO-management problem rather than a pure model-optimization problem. ...

November 1, 2024 · Publish Assistant

OSDI 2024 Digest

11 papers selected. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al. TL;DR — Separates the compute-heavy prefill phase from the memory-bound decoding phase onto different GPU pools, eliminating head-of-line blocking and significantly improving LLM serving throughput. Why notable — Became one of the most influential LLM systems papers of 2024; the prefill–decode disaggregation insight is now widely adopted in production inference stacks (vLLM, SGLang, etc.). ...

July 10, 2024 · Publish Assistant

CCGrid 2024 Digest

10 papers selected. Fair, Efficient Multi-Resource Scheduling for Stateless Serverless Functions with Anubis Amit Samanta 0001, Ryan Stutsman TL;DR — Anubis introduces a fair, multi-resource scheduler for stateless serverless functions that achieves efficiency without sacrificing isolation between tenants. Why notable — Fairness in serverless resource allocation is an open problem as functions compete for heterogeneous resources (CPU, memory, I/O); Anubis provides a concrete, deployable answer. The work directly addresses a gap in production FaaS platforms where existing schedulers optimize for throughput but ignore per-tenant equity. ...

May 6, 2024 · Publish Assistant

ATC 2024 Digest

12 papers selected. FetchBPF: Customizable Prefetching Policies in Linux with eBPF Xuechun Cao, Shaurya Patel, Soo-Yee Lim, Xueyuan Han et al. TL;DR — Extends eBPF into the page-fault / prefetch path, giving user-space programs a safe, low-overhead hook to install custom hardware-prefetch policies without kernel modifications. Fast (Trapless) Kernel Probes Everywhere Jinghao Jia, Michael V. Le, Salman Ahmed 0001, Dan Williams 0001 et al. TL;DR — Eliminates the trap-based overhead of kprobes by using binary rewriting to instrument kernel functions at near-zero cost, enabling always-on production tracing. ...

January 1, 2024 · Publish Assistant

EuroSys 2024 Digest

13 papers selected. Pronghorn: Effective Checkpoint Orchestration for Serverless Hot-Starts Sumer Kohli, Shreyas Kharbanda, Rodrigo Bruno, João Carreira et al. TL;DR — Demonstrates how carefully orchestrated checkpointing can eliminate cold-start latency in serverless runtimes, achieving near-instant hot-starts with negligible overhead. Serialization/Deserialization-free State Transfer in Serverless Workflows Fangming Lu, Xingda Wei, Zhuobin Huang, Rong Chen 0001 et al. TL;DR — Eliminates the dominant serialization cost in serverless function chaining by enabling direct in-memory state passing, yielding large end-to-end latency reductions. ...

January 1, 2024 · Publish Assistant

FGCS 2024 Digest

12 papers selected. Quantum-centric supercomputing for materials science: A perspective on challenges and future directions Yuri Alexeev, Maximilian Amsler, Marco Antonio Barroca, Sanzio Bassini et al. TL;DR — A comprehensive roadmap from IBM, national labs, and universities identifying key algorithmic, software, and hardware challenges for using quantum processors alongside classical HPC to advance materials science simulations. Why notable — Essential reading for any researcher planning quantum-classical hybrid workflows, covering the full stack from error mitigation to application mapping at scale. ...

January 1, 2024 · Publish Assistant

HPDC 2024 Digest

12 papers selected. Efficient all-to-all Collective Communication Schedules for Direct-connect Topologies Prithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal et al. TL;DR — Derives near-optimal all-to-all collective communication schedules for direct-connect HPC topologies, directly improving bandwidth utilization in large-scale distributed systems. Reinforcement Learning-based Adaptive Mitigation of Uncorrected DRAM Errors in the Field Isaac Boixaderas, Sergi Moré, Javier Bartolome, David Vicente et al. TL;DR — Applies reinforcement learning to dynamically mitigate uncorrected DRAM errors at production HPC scale, improving system reliability without sacrificing performance. ...

January 1, 2024 · Publish Assistant

IC 2024 Digest

10 papers selected. Revisiting Edge AI: Opportunities and Challenges Tobias Meuser, Lauri Lovén, Monowar Bhuyan, Shishir G. Patil et al. TL;DR — A multi-author position paper that revisits the state of edge AI, cataloguing deployment barriers and open research problems across hardware, networking, and software layers. Why notable — Brings together 19 leading researchers to synthesize the field’s most pressing edge AI challenges, making it an authoritative reference for practitioners and researchers planning edge deployments. ...

January 1, 2024 · Publish Assistant

IPDPS 2024 Digest

14 papers selected. Low-Depth Spatial Tree Algorithms Yves Baumann, Tal Ben-Nun, Maciej Besta, Lukas Gianinazzi et al. TL;DR — Introduces parallel spatial-tree algorithms with provably low depth, advancing the theory of work-efficient parallel data structures for geometric workloads. Alternative Basis Matrix Multiplication is Fast and Stable Oded Schwartz, Sivan Toledo, Noa Vaknin, Gal Wiernik TL;DR — Demonstrates that alternative-basis matrix multiplication achieves both practical speed and numerical stability, challenging the conventional trade-off between the two. ...

January 1, 2024 · Publish Assistant

JPDC 2024 Digest

12 papers selected. Read/write fence-free work-stealing with multiplicity Armando Castañeda, Miguel Piña TL;DR — Presents a work-stealing deque algorithm that eliminates read/write memory fences while tolerating multiplicity, achieving provably correct concurrent access without costly barriers. Why notable — Advances the theoretical foundations of lock-free scheduler data structures by decoupling correctness from fence instructions, directly impacting runtime system design. Reliable communication in dynamic networks with locally bounded byzantine faults Silvia Bonomi, Giovanni Farina, Sébastien Tixeuil ...

January 1, 2024 · Publish Assistant

MobiSys 2024 Digest

13 papers selected. WAIS: Leveraging WiFi for Resource-Efficient SLAM Aditya Arun 0002, William Hunter, Roshan Sai Ayyalasomayajula, Dinesh Bharadia TL;DR — Demonstrates that commodity WiFi signals can replace LiDAR for simultaneous localization and mapping, dramatically cutting the resource cost of robot/AR navigation. UWB-Fi: Pushing Wi-Fi towards Ultra-wideband for Fine-Granularity Sensing Xin Li 0070, Hongbo Wang, Zhe Chen 0015, Zhiping Jiang et al. TL;DR — Extends standard Wi-Fi to UWB-class sensing resolution without hardware changes, enabling centimeter-level gesture and motion detection on existing infrastructure. ...

January 1, 2024 · Publish Assistant

NSDI 2024 Digest

13 papers selected. MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUs Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang et al. TL;DR — ByteDance’s full production account of training LLMs at 10,000+ GPUs, with novel co-design of the network stack, fault tolerance, and collective communication to sustain near-linear scaling. Harmony: A Congestion-free Datacenter Architecture Saksham Agarwal, Qizhe Cai, Rachit Agarwal 0001, David B. Shmoys et al. ...

January 1, 2024 · Publish Assistant

SC 2024 Digest

15 papers selected. Pushing the Limit of Quantum Mechanical Simulation to the Raman Spectra of a Biological System with 100 Million Atoms Honghui Shang, Ying Liu 0055, Zhikun Wu, Zhenchuan Chen et al. TL;DR — Gordon Bell-class result scaling ab initio Raman spectroscopy to 100 million atoms, pushing quantum-chemical simulation well beyond prior limits. Breaking the Molecular Dynamics Timescale Barrier Using a Wafer-Scale System Kylee Santos, Stan G. Moore, Tomas Oppelstrup, Amirali Sharifian et al. ...

January 1, 2024 · Publish Assistant

SEC 2024 Digest

12 papers selected. EdgeCore: Resource Dependency-Aware Multi-Tenant Orchestration for Mobile Edge Clouds Amran Haroon TL;DR — Introduces a multi-tenant edge orchestration system that captures resource dependencies across co-located workloads, demonstrating significant improvements in task completion latency and resource utilization. Righteous: Automatic Right-Sizing for Complex Edge Deployments Aniruddha Rakshit TL;DR — Presents an automated right-sizing framework for edge deployments that dynamically adjusts resource allocations to match workload demands without manual intervention. ...

January 1, 2024 · Publish Assistant

TC 2024 Digest

12 papers selected. Achieving DRAM-Like PCM by Trading Off Capacity for Latency Irina Alam, Puneet Gupta 0001 TL;DR — Proposes a capacity-for-latency trade-off in Phase Change Memory to match DRAM-level access latency without specialized process changes. Why notable — Offers a practical path to deploying PCM as a DRAM alternative, directly addressing the latency gap that has blocked PCM adoption in main-memory systems. A High-Performance, Energy-Efficient Modular DMA Engine Architecture Thomas Benz, Michael Rogenmoser, Paul Scheffler, Samuel Riedel et al. ...

January 1, 2024 · Publish Assistant

TCC 2024 Digest

12 papers selected. FaaSCtrl: A Comprehensive-Latency Controller for Serverless Platforms Abhisek Panda, Smruti R. Sarangi TL;DR — FaaSCtrl is a feedback-control system for serverless platforms that jointly manages cold-start, queuing, and execution latency to meet end-to-end SLOs. Why notable — One of the few serverless controllers that addresses all three latency components together, providing a principled alternative to ad-hoc autoscaling heuristics. FUSIONIZE++: Improving Serverless Application Performance Using Dynamic Task Inlining and Infrastructure Optimization Trever Schirmer, Joel Scheuner, Tobias Pfandzelter, David Bermbach ...

January 1, 2024 · Publish Assistant

TOCS 2024 Digest

8 papers selected. PMAlloc: A Holistic Approach to Improving Persistent Memory Allocation Zheng Dang, Shuibing He, Xuechen Zhang, Peiyi Hong et al. TL;DR — PMAlloc redesigns persistent memory allocation end-to-end, co-optimizing the allocator’s data structures, concurrency, and crash consistency to dramatically reduce allocation overhead. Why notable — Persistent memory is still poorly understood at the allocator level; this paper offers a rare holistic treatment that will inform future PM software stacks. ...

January 1, 2024 · Publish Assistant

TPDS 2024 Digest

12 papers selected. Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al. TL;DR — An active-learning framework automatically diagnoses runtime performance anomalies in production HPC systems by querying targeted job profiles to minimize labeling effort. Why notable — It bridges ML-based anomaly detection and operational HPC monitoring, demonstrating scalable root-cause identification on real supercomputer workloads. AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost Jinfan Chen, Shigang Li 0002, Ran Guo, Jinhui Yuan et al. ...

January 1, 2024 · Publish Assistant