Balancing Repair Bandwidth and Sub-Packetization in Erasure-Coded Storage via Elastic Transformation
Kaicheng Tang, Keyun Cheng, Helen H. W. Chan, Xiaolu Li, Patrick P. C. Lee, Yuchong Hu, Jie Li, Ting-Yi Wu
摘要
Erasure coding provides high fault-tolerant storage with significantly low redundancy overhead, at the expense of high repair bandwidth. While there exist access-optimal codes that theoretically minimize both the repair bandwidth and the amount of disk reads, they also incur a high sub-packetization level, thereby leading to non-sequential I/Os and degrading repair performance. We propose elastic transformation, a framework that transforms any base code into a new code with smaller repair bandwidth for all or a subset of nodes, such that it can be configured with a wide range of sub-packetization levels to limit the non-sequential I/O overhead. We prove the fault tolerance of elastic transformation and model numerically the repair performance with respect to a sub-packetization level. We further prototype and evaluate elastic transformation atop HDFS, and show how it reduces the single-block repair time of the base codes and access-optimal codes in a real network setting.
• We show how our elastic transformation reduces the repair bandwidth of various erasure codes, including RS codes [32], Azure's LRC [13], Hitchhiker [31], and HashTag [18]. • On the theoretical side, we prove the fault tolerance of elastic transformation, model the lower bound of repair bandwidth for a given sub-packetization level, and model the repair time subject to the bandwidth and I/O conditions. • Existing code transformation solutions [11], [12], [19], [20]
only focus on theoretical analysis, but do not consider empirical evaluation. To fill this void, we prototype elastic transformation based on OpenEC [23] and evaluate it atop Hadoop 3.0.0 HDFS [1]. Experiments on our prototype, called OpenEC-ET, show that it reduces the repair time of the base RS codes (without transformation) by up to 56.3% in low-bandwidth settings and reduces the repair time of the access-optimal MSR codes by up to 51.4% in highbandwidth settings. The source code of OpenEC-ET is at: http://adslab.cse.cuhk.edu.hk/software/openec-et.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Revisiting Network Coding for Warm Blob StorageChuang Gan, Yuchong Hu, Leyan Zhao, Xin Zhao 等FAST 2025 · 被引用 9 次
- LESS is More for I/O-Efficient Repairs in Erasure-Coded StorageKeyun Cheng, Guodong Li, Xiaolu Li, Sihuang Hu 等FAST 2026 · 被引用 4 次
- Leveled Product Codes for Optimal Block Repairs in Geo-distributed Storage SystemsSi Wu, Guantian Lin, Patrick P. C. Lee, Yinlong XuINFOCOM 2025 · 被引用 2 次
- WiseCode: Breaking the Scalability Barriers of Wide-Stripe Vector CodesSijie Cai, Guangyan Zhang, Xiao NiuOSDI 2026
它引用的顶会 Paper1
相关 Paper
- ParaRC: Embracing Sub-Packetization for Repair Parallelization in MSR-Coded StorageXiaolu Li, Keyun Cheng, Kaicheng Tang, Patrick P. C. Lee 等FAST 2023
- ChameleonEC: Exploiting Tunability of Erasure Coding for Low-Interference RepairYuhui Cai, Shiyao Lin, Zhirong Shen, Jiahui Yang 等HPCA 2025 · 被引用 5 次
- Optimal Data Placement for Stripe Merging in Locally Repairable CodesSi Wu, Qingpeng Du, Patrick P. C. Lee, Yongkun Li 等INFOCOM 2022 · 被引用 23 次
- Boosting Full-Node Repair in Erasure-Coded StorageShiyao Lin, Guowen Gong, Zhirong Shen, Patrick P. C. Lee 等USENIX ATC 2021 · 被引用 33 次
- PDL: A Data Layout towards Fast Failure Recovery for Erasure-coded Distributed Storage SystemsLiangliang Xu, Min Lv, Zhipeng Li, Cheng Li 等INFOCOM 2020 · 被引用 13 次
