The Dilemma between Deduplication and Locality: Can Both be Achieved?
Xiangyu Zou, Jingsong Yuan, Philip Shilane, Wen Xia, Haijun Zhang, Xuan Wang
摘要
Data deduplication is widely used to reduce the size of backup workloads, but it has the known disadvantage of causing poor data locality, also referred to as the fragmentation problem, which leads to poor restore and garbage collection (GC) performance. Current research has considered writing duplicates to maintain locality (e.g. rewriting) or caching data in memory or SSD, but fragmentation continues to hurt restore and GC performance. Investigating the locality issue, we observed that most duplicate chunks in a backup are directly from its previous backup. We therefore propose a novel management-friendly deduplication framework, called MFDedup, that maintains the locality of backup workloads by using a data classification approach to generate an optimal data layout. Specifically, we use two key techniques: Neighbor-Duplicate-Focus indexing (NDF) and Across-Version-Aware Reorganization scheme (AVAR), to perform duplicate detection against a previous backup and then rearrange chunks with an offline and iterative algorithm into a compact, sequential layout that nearly eliminates random I/O during restoration. Evaluation results with four backup datasets demonstrates that, compared with state-of-the-art techniques, MFDedup achieves deduplication ratios that are 1.12x to 2.19x higher and restore throughputs that are 2.63x to 11.64x faster due to the optimal data layout we achieve. While the rearranging stage introduces overheads, it is more than offset by a nearly-zero overhead GC process. Moreover, the NDF index only requires indexes for two backup versions, while the traditional index grows with the number of versions retained.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Secure and Lightweight Deduplicated Storage via Shielded Deduplication-Before-EncryptionZuoru Yang, Jingwei Li, Patrick P. C. LeeUSENIX ATC 2022 · 被引用 46 次
- Building a High-performance Fine-grained Deduplication Framework for Backup Storage with High Deduplication RatioXiangyu Zou, Wen Xia, Philip Shilane, Haijun Zhang 等USENIX ATC 2022 · 被引用 36 次
- TiDedup: A New Distributed Deduplication Architecture for CephMyoungwon Oh, Sungmin Lee, Samuel Just, Youngjin Yu 等USENIX ATC 2023 · 被引用 23 次
- LoopDelta: Embedding Locality-aware Opportunistic Delta Compression in Inline Deduplication for Highly Efficient Data ReductionYucheng Zhang, Hong Jiang, Dan Feng, Nan Jiang 等USENIX ATC 2023 · 被引用 13 次
- Don't Maintain Twice, It's Alright: Merged Metadata Management in Deduplication File System with GogetaFSYanqi Pan, Wen Xia, Erci Xu, Hao Huang 等FAST 2025 · 被引用 6 次
它引用的顶会 Paper1
相关 Paper
- Garbage Collection Does Not Only Collect Garbage: Piggybacking-Style Defragmentation for Deduplicated Backup StorageDingbang Liu, Xiangyu Zou, Tao Lu, Philip Shilane 等EuroSys 2025 · 被引用 1 次
- Once Rolling Hashing is Enough: Exploiting Rolling Hash Reuse in Delta CompressionHaoliang Tan, Wenhao Ou, Xiangyu Zou, Cai Deng 等EuroSys 2026 · 被引用 1 次
- FinerDedup: Sifting Fingerprints for Efficient Data Deduplication on Mobile DevicesXianzhang Chen, Xingjie Zhou, Wei Li, Xi Yu 等DAC 2024 · 被引用 2 次
- CDCache: Space-Efficient Flash Caching via Compression-before-DeduplicationHengying Xiao, Jingwei Li, Yanjing Ren, Ruijin Wang 等INFOCOM 2024 · 被引用 1 次
- Physical vs. Logical Indexing with IDEA: Inverted Deduplication-Aware IndexAsaf Levi, Philip Shilane, Sarai Sheinvald, Gala YadgarFAST 2024 · 被引用 7 次
