IOCost: block IO control for containers in datacenters
Tejun Heo, Dan Schatzberg, Andrew Newell, Song Liu, Saravanan Dhakshinamurthy, Iyswarya Narayanan, Josef Bacik, Chris Mason, Chunqiang Tang, Dimitrios Skarlatos
摘要
Resource isolation is a fundamental requirement in datacenter environments. However, our production experience in Meta’s large-scale datacenters shows that existing IO control mechanisms for block storage are inadequate in containerized environments. IO control needs to provide proportional resources to containers while taking into account the hardware heterogeneity of storage devices and the idiosyncrasies of the workloads deployed in datacenters. The speed of modern SSDs requires IO control to execute with low-overheads. Furthermore, IO control should strive for work conservation, take into account the interactions with the memory management subsystem, and avoid priority inversions that lead to isolation failures. To address these challenges, this paper presents IOCost, an IO control solution that is designed for containerized environments and provides scalable, work-conserving, and low-overhead IO control for heterogeneous storage devices and diverse workloads in datacenters. IOCost performs offline profiling to build a device model and uses it to estimate device occupancy of each IO request. To minimize runtime overhead, it separates IO control into a fast per-IO issue path and a slower periodic planning path. A novel work-conserving budget donation algorithm enables containers to dynamically share unused budget. We have deployed IOCost across the entirety of Meta’s datacenters comprised of millions of ma- chines, upstreamed IOCost to the Linux kernel, and open-sourced our device-profiling tools. IOCost has been running in production for two years, providing IO control for Meta’s fleet. We describe the design of IOCost and share our experience deploying it at scale.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Disaggregated RAID Storage in Modern DatacentersJunyi Shu, Ruidong Zhu, Yun Ma, Gang Huang 等ASPLOS 2023 · 被引用 18 次
- Burstable Cloud Block Storage with Data Processing UnitsJunyi Shu, Kun Qian, Ennan Zhai, Xuanzhe Liu 等OSDI 2024 · 被引用 17 次
- Xkernel: Principled Performance Tunability of Operating System KernelsZhongjie Chen, Wentao Zhang, Yulong Tang, Ran Shu 等OSDI 2026 · 被引用 2 次
它引用的顶会 Paper4
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor 等OSDI 2020 · 被引用 107 次
- RAS: Continuously Optimized Region-Wide Datacenter Resource AllocationAndrew Newell, Dimitrios Skarlatos, Jingyuan Fan, Pavan Kumar 等SOSP 2021 · 被引用 19 次
- BabelFish: Fusing Address Translations for ContainersDimitrios Skarlatos, Umur Darbaz, Bhargava Gopireddy, Nam Sung Kim 等ISCA 2020 · 被引用 18 次
- Draco: Architectural and Operating System Support for System Call SecurityDimitrios Skarlatos, Qingrong Chen, Jianyan Chen, Tianyin Xu 等MICRO 2020 · 被引用 17 次
相关 Paper
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang 等ASPLOS 2022 · 被引用 103 次
- i-NVMe: Isolated NVMe over TCP for a Containerized EnvironmentSeongho Lee, Ikjun Yeom, Younghoon KimINFOCOM 2023
- Harvesting Spare CPU Resources in Container SystemsAdam Hall, Anirudh Sarma, Esha Choukse, Umakishore Ramachandran 等NSDI 2026 · 被引用 2 次
- D2FQ: Device-Direct Fair Queueing for NVMe SSDsJiwon Woo, Minwoo Ahn, Gyusun Lee, Jinkyu JeongFAST 2021 · 被引用 45 次
- Fair Will Go On: A Collaboration-Aware Fairness Scheme for NVMe SSD in Cloud Storage SystemYang Zhou, Fang Wang, Zhan Shi, Dan Feng 等DAC 2023 · 被引用 7 次
