MAPX: Controlled Data Migration in the Expansion of Decentralized Object-Based Storage Systems
Li Wang, Yiming Zhang, Jiawei Xu, Guangtao Xue
Abstract
Data placement is critical for the scalability of decentralized object-based storage systems. The state-of-the-art CRUSH placement method is a decentralized algorithm that deterministically places object replicas onto storage devices without relying on a central directory. While enjoying the benefits of decentralization such as high scalability, robustness, and performance, CRUSH-based storage systems suffer from uncontrolled data migration when expanding the clusters, which will cause significant performance degradation when the expansion is nontrivial.
This paper presents MAPX, a novel extension to CRUSH that uses an extra time-dimension mapping (from object creation times to cluster expansion times) for controlled data migration in cluster expansions. Each expansion is viewed as a new layer of the CRUSH map represented by a virtual node beneath the CRUSH root. MAPX controls the mapping from objects onto layers by manipulating the timestamps of the intermediate placement groups (PGs). MAPX is applicable to a large variety of object-based storage scenarios where object timestamps can be maintained as higher-level metadata. For example, we apply MAPX to Ceph-RBD by extending the RBD metadata structure to maintain and retrieve approximate object creation times at the granularity of expansions layers. Experimental results show that the MAPX-based migration-free system outperforms the CRUSH-based system (which is busy in migrating objects after expansions) by up to 4.25× in the tail latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d6a19c9-e6dc-4fd6-a9c9-de5f0c3b5f4bCited by top-tier papers5
- Facebook's Tectonic Filesystem: Efficiency from ExascaleSatadru Pan, Theano Stavrinos, Yunqiao Zhang, Atul Sikaria et al.FAST 2021 · 110 citations
- DEPART: Replica Decoupling for Distributed Key-Value StorageQiang Zhang, Yongkun Li, Patrick P. C. Lee, Yinlong Xu et al.FAST 2022 · 12 citations
- MapperX: Adaptive Metadata Maintenance for Fast Crash Recovery of DM-Cache Based Hybrid Storage DevicesLujia Yin, Li Wang, Yiming Zhang, Yuxing PengUSENIX ATC 2021 · 7 citations
- Provably Good Randomized Strategies for Data Placement in Distributed Key-Value StoresZhe Wang, Jinhao Zhao, Kunal Agrawal, He Liu et al.PPoPP 2023 · 3 citations
- Cheetah: Metadata Aggregation for Fast Object Storage without Distributed OrderingYiming Zhang, Li Wang, Shengyun Liu, Shun Gai et al.EuroSys 2025
Related papers
- The what, The from, and The to: The Migration Games in Deduplicated SystemsRoei Kisous, Ariel Kolikant, Abhinav Duggal, Sarai Sheinvald et al.FAST 2022 · 12 citations
- TiDedup: A New Distributed Deduplication Architecture for CephMyoungwon Oh, Sungmin Lee, Samuel Just, Youngjin Yu et al.USENIX ATC 2023 · 23 citations
- Separating Data via Block Invalidation Time Inference for Write Amplification Reduction in Log-Structured StorageQiuping Wang, Jinhong Li, Patrick P. C. Lee, Tao Ouyang et al.FAST 2022 · 56 citations
- GeoLayer: Towards Low-Latency and Cost-Efficient Geo-Distributed Graph Stores with Layered GraphFeng Yao, Xiaokang Yang, Shufeng Gong, Song Yu et al.ICDE 2026 · 1 citation
- Migration-Free Elastic Storage of Time Series in Apache IoTDBRongzhao Chen, Xiangpeng Hu, Xiangdong Huang, Chen Wang et al.VLDB 2025
