Scheduling Cloud Block Storage Proactively and Reactively with Omar
Xinqi Chen, Weidong Zhang, Zhongyu Wang, Erci Xu, Xiaolu Zhang, Dong Wu, Junping Wu, Haonan Wu, Ruiming Lu, Yaheng Song, Chaolei Hu, Lijun Ding
Abstract
We explore the performance improvements for a production Elastic Block Storage (EBS) service at Alibaba, which serves millions of users and virtual disks daily. Despite various load balancing measures, Alibaba EBS still faces frequent performance variations or even occasional Service Level Objective (SLO) violations. Our trace analysis reveals that the existing scheduling mechanism in Alibaba EBS, which adopts a common practice of focusing on load balancing alone, is insufficient for mitigating tail latency due to negligence of correlated I/O workloads, reliance on fixed scheduling frequencies, and insufficient input metrics for scheduling decisions. We present Omar, a novel hybrid proactive and reactive scheduling framework that enables dynamic, finegrained scheduling. Evaluation on production-scale testbeds shows that Omar reduces 64 KiB tail latency by up to 52.2% and the number of scheduling operations by 44.6% compared to the baseline scheduler, while incurring limited overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9929815c-76ee-410d-a126-45e4473d60b2Builds on20
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM ServingBingyang Wu, Ruidong Zhu, Zili Zhang, Peng Sun et al.OSDI 2024 · 79 citations
- Take it to the limit: peak prediction-driven resource overcommitment in datacentersNoman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin et al.EuroSys 2021 · 60 citations
- Separating Data via Block Invalidation Time Inference for Write Amplification Reduction in Log-Structured StorageQiuping Wang, Jinhong Li, Patrick P. C. Lee, Tao Ouyang et al.FAST 2022 · 56 citations
Related papers
- Come Hell or Still Water: Alleviating Tail Latency in Cloud Block StoreChaolei Hu, Kun Qian, Erci Xu, Yifan Shen et al.NSDI 2026
- Hey Hey, My My, Skewness Is Here to Stay: Challenges and Opportunities in Cloud Block Store TrafficHaonan Wu, Erci Xu, Ligang Wang, Yuandong Hong et al.EuroSys 2025 · 2 citations
- Burstable Cloud Block Storage with Data Processing UnitsJunyi Shu, Kun Qian, Ennan Zhai, Xuanzhe Liu et al.OSDI 2024 · 17 citations
- How Soon is Now? Preloading Images for Virtual Disks with ThinkAheadXinqi Chen, Yu Zhang, Erci Xu, Changhong Wang et al.FAST 2026 · 1 citation
- Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud PlatformsChengzhi Lu, Huanle Xu, Kejiang Ye, Guoyao Xu et al.EuroSys 2023 · 34 citations
