PACEMAKER: Avoiding HeART attacks in storage clusters with disk-adaptive redundancy
Saurabh Kadekodi, Francisco Maturana, Suhas Jayaram Subramanya, Juncheng Yang, K. V. Rashmi, Gregory R. Ganger
Abstract
Data redundancy provides resilience in large-scale storage clusters, but imposes significant cost overhead. Substantial space-savings can be realized by tuning redundancy schemes to observed disk failure rates. However, prior design proposals for such tuning are unusable in real-world clusters, because the IO load of transitions between schemes overwhelms the storage infrastructure (termed transition overload). This paper analyzes traces for millions of disks from production systems at Google, NetApp, and Backblaze to expose and understand transition overload as a roadblock to diskadaptive redundancy: transition IO under existing approaches can consume 100% cluster IO continuously for several weeks. Building on the insights drawn, we present PACEMAKER, a low-overhead disk-adaptive redundancy orchestrator. PACE-MAKER mitigates transition overload by (1) proactively organizing data layouts to make future transitions efficient, and (2) initiating transitions proactively in a manner that avoids urgency while not compromising on space-savings. Evaluation of PACEMAKER with traces from four large (110K-450K disks) production clusters show that the transition IO requirement decreases to never needing more than 5% cluster IO bandwidth (0.2-0.4% on average). PACEMAKER achieves this while providing overall space-savings of 14-20% and never leaving data under-protected. We also describe and experiment with an integration of PACEMAKER into HDFS. 1 AFR describes the expected fraction of disks that experience failure in a typical year.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9eaee442-85f4-49da-9385-a6d054a2cb62Cited by top-tier papers12
- Practical Design Considerations for Wide Locally Recoverable Codes (LRCs)Saurabh Kadekodi, Shashwat Silas, David Clausen, Arif MerchantFAST 2023 · 55 citations
- Optimal Data Placement for Stripe Merging in Locally Repairable CodesSi Wu, Qingpeng Du, Patrick P. C. Lee, Yongkun Li et al.INFOCOM 2022 · 23 citations
- Tiger: Disk-Adaptive Redundancy Without Placement RestrictionsSaurabh Kadekodi, Francisco Maturana, Sanjith Athlur, Arif Merchant et al.OSDI 2022 · 20 citations
- Disaggregated RAID Storage in Modern DatacentersJunyi Shu, Ruidong Zhu, Yun Ma, Gang Huang et al.ASPLOS 2023 · 18 citations
- ELECT: Enabling Erasure Coding Tiering for LSM-tree-based StorageYanjing Ren, Yuanming Ren, Xiaolu Li, Yuchong Hu et al.FAST 2024 · 18 citations
Related papers
- Morph: Efficient File-Lifetime Redundancy Management for Cluster File SystemsTimothy Kim, Sanjith Athlur, Saurabh Kadekodi, Francisco Maturana et al.SOSP 2024 · 4 citations
- Thesios: Synthesizing Accurate Counterfactual I/O Traces from I/O SamplesPhitchaya Mangpo Phothilimthana, Saurabh Kadekodi, Soroush Ghodrati, Selene Moon et al.ASPLOS 2024 · 6 citations
- PASS: A Power Adaptive Storage ServerDedong Xie, Theano Stavrinos, Jonggyu Park, Simon Peter et al.EuroSys 2026
- BCW: Buffer-Controlled Writes to HDDs for SSD-HDD Hybrid Storage ServerShucheng Wang, Ziyi Lu, Qiang Cao, Hong Jiang et al.FAST 2020 · 37 citations
- Take it to the limit: peak prediction-driven resource overcommitment in datacentersNoman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin et al.EuroSys 2021 · 60 citations
