All checks were successful
Build and deploy static pages / build-and-push (push) Successful in 19s
144 lines
5.7 KiB
YAML
144 lines
5.7 KiB
YAML
venue: SoCC
|
|
year: 2025
|
|
papers:
|
|
- title: 'From Bottleneck to Breakthrough: Optimizing Scheduling for Hyperscale Containerized
|
|
Clusters'
|
|
authors:
|
|
- Bing Li
|
|
- Yuquan Ren
|
|
- Xinyi Song
|
|
- Zhilei Liu
|
|
- Cong Xu
|
|
- Jingyuan Zhang
|
|
- Caixue Lin
|
|
- Wu Xiang
|
|
- Rui Shi
|
|
reason: "Documents production-scale scheduling improvements at a hyperscale cloud provider, demonstrating how targeted optimizations reduce scheduling tail latency and increase cluster utilization in real containerized workloads."
|
|
- title: 'CPU-Limits kill Performance: Time to rethink Resource Control'
|
|
authors:
|
|
- Chirag C. Shetty
|
|
- Sarthak Chakraborty
|
|
- Hubertus Franke
|
|
- Larisa Shwartz
|
|
- Chandra Narayanaswami
|
|
- Indranil Gupta
|
|
- Saurabh Jha
|
|
reason: "Challenges the conventional use of CPU cgroup limits in cloud environments, showing through production evidence that CFS bandwidth throttling degrades application QoS and proposing a rethink of resource control abstractions."
|
|
- title: Rethinking Tiered Memory Management in Cloud Data Centers
|
|
authors:
|
|
- Tong Xing 0002
|
|
- Jiaxun Yang
|
|
- Javier Picorel
|
|
- Antonio Barbalace
|
|
reason: "Proposes a novel tiered memory management framework for cloud data centers that improves performance by rethinking the placement and migration policies across DRAM and CXL/NVM tiers."
|
|
- title: Cost-Efficient Cloud Infrastructure with Hugepage-aware Memory Deduplication
|
|
authors:
|
|
- Ruizhe Huang
|
|
- Xinyu Wang 0043
|
|
- Zhida An
|
|
- Hanwen Lei
|
|
- Peng Jiang 0007
|
|
- Ziqi Zhang
|
|
- Ding Li 0001
|
|
- Yao Guo 0001
|
|
- Xiangqun Chen
|
|
- Yuntao Liu
|
|
- Kang Zhou
|
|
- Yuxin Ren 0001
|
|
- Ning Jia 0004
|
|
- Xinwei Hu
|
|
reason: "Deploys hugepage-aware memory deduplication in a large production cloud, achieving significant memory savings without the performance regressions that plague conventional THP-based deduplication."
|
|
- title: 'ALAP: Intent-Based Serverless Computing via Delayed Decision-Making'
|
|
authors:
|
|
- Prasoon Sinha
|
|
- Kostis Kaffes
|
|
- Neeraja J. Yadwadkar
|
|
reason: "Introduces an intent-based programming model for serverless that defers scheduling decisions until runtime context is available, improving resource efficiency and SLO attainment over eager placement strategies."
|
|
- title: 'Hydra: Virtualized Multi-Language Runtime for High-Density Serverless Platforms'
|
|
authors:
|
|
- Serhii Ivanenko
|
|
- Vasyl Lanko
|
|
- Rudi Horn
|
|
- Vojin Jovanovic
|
|
- Rodrigo Bruno
|
|
reason: "Presents a virtualized runtime that multiplexes multiple language environments within a single sandbox, enabling higher function density and faster cold starts on serverless platforms."
|
|
- title: 'Serverless Elasticsearch: the Architecture Transformation from Stateful to
|
|
Stateless'
|
|
authors:
|
|
- Iraklis Psaroudakis
|
|
- Pooya Salehi
|
|
- Jason Bryan
|
|
- Francisco Fernández Castaño
|
|
- Brendan Cully
|
|
- Ankita Kumar
|
|
- Henning Andersen
|
|
- Thomas Repantis
|
|
reason: "Describes Elastic's production migration of Elasticsearch to a serverless, stateless architecture, sharing engineering lessons on decoupling compute from state at cloud scale."
|
|
- title: 'DFUSE: Strongly Consistent Write-Back Kernel Caching for Distributed Userspace
|
|
File Systems'
|
|
authors:
|
|
- Haoyu Li
|
|
- Jingkai Fu
|
|
- Qing Li 0002
|
|
- Windsor Hsu
|
|
- Asaf Cidon
|
|
reason: "Closes a long-standing gap in FUSE-based distributed file systems by enabling strongly consistent write-back caching in the kernel, significantly improving throughput without sacrificing correctness."
|
|
- title: Accelerating Distributed Filesystem Metadata Service via Decoupling Directory
|
|
Semantics from Metadata Indexing
|
|
authors:
|
|
- Wenhao Lv
|
|
- Hao Guo
|
|
- Qing Wang 0031
|
|
- Youyou Lu
|
|
- Jiwu Shu
|
|
reason: "Achieves scalable distributed filesystem metadata by separating directory namespace semantics from the underlying index structure, reducing contention and improving throughput for large-scale cloud storage."
|
|
- title: 'Valet: Efficient Data Placement on Modern SSDs'
|
|
authors:
|
|
- Devashish R. Purandare
|
|
- Peter Alvaro
|
|
- Avani Wildani
|
|
- Darrell D. E. Long
|
|
- Ethan L. Miller
|
|
reason: "Exploits fine-grained internal SSD geometry to make smarter data placement decisions, yielding measurable I/O performance gains without changes to the host storage stack."
|
|
- title: 'Understanding Diffusion Model Serving in Production: A Top-Down Analysis
|
|
of Workload, Scheduling, and Resource Efficiency'
|
|
authors:
|
|
- Yanying Lin
|
|
- Shuaipeng Wu
|
|
- Shutian Luo
|
|
- Hong Xu 0001
|
|
- Haiying Shen
|
|
- Chong Ma
|
|
- Min Shen
|
|
- Le Chen
|
|
- Chengzhong Xu 0001
|
|
- Lin Qu
|
|
- Kejiang Ye
|
|
reason: "Provides the first comprehensive production characterization of diffusion model inference workloads, revealing unique scheduling and resource efficiency challenges distinct from LLM serving."
|
|
- title: 'ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable
|
|
Multimodal Model Serving'
|
|
authors:
|
|
- Haoran Qiu
|
|
- Anish Biswas
|
|
- Zihan Zhao
|
|
- Jayashree Mohan
|
|
- Alind Khare
|
|
- Esha Choukse
|
|
- Íñigo Goiri
|
|
- Zeyu Zhang 0005
|
|
- Haiying Shen
|
|
- Chetan Bansal
|
|
- Ramachandran Ramjee
|
|
- Rodrigo Fonseca
|
|
reason: "Disaggregates compute resources per modality and pipeline stage for multimodal inference, with a Microsoft production deployment showing improved GPU utilization and latency over monolithic serving."
|
|
- title: 'THORN-ML: Transparent Hardware Offloaded Resilient Networks for RDMA based
|
|
Distributed ML Workloads'
|
|
authors:
|
|
- Maziyar Nazari
|
|
- Daniel Noland
|
|
- Giulio Sidoretti
|
|
- Erika Hunhoff
|
|
- Tamara Silbergleit Lehman
|
|
- Eric Keller
|
|
reason: "Offloads RDMA fault detection and recovery to programmable network hardware, making distributed ML training resilient to network failures without modifying the training framework or incurring software overhead."
|