Odess: Speeding up Resemblance Detection for Redundancy Elimination by Fast Content-Defined Sampling
Xiangyu Zou, Cai Deng, Wen Xia, Philip Shilane, Haoliang Tan, Haijun Zhang, Xuan Wang
Abstract
Multiple data reduction techniques have been investigated to lower storage costs for a wide variety of customers. In this work, we focus on similarity-based delta compression, which calculates and stores the difference of very similar, but non-duplicate, chunks in storage systems. Delta compression is often implemented along with deduplication and has been shown to achieve a much higher compression ratio. Currently, the N-Transform method is the most popular and widely-used approach to generate features for data content (e.g. chunks) to detect similar candidates (and then apply delta compression). For delta compression systems, though, the throughput of N-Transform is often the bottleneck. Finesse is a high throughput variant of N-Transform, but it suffers from lower detection accuracy and compression ratio. The computation overhead of N-Transform consists of two parts: calculating the rolling hash across data and applying time-consuming transforms on each hash. In this work, we propose Odess, a fast resemblance detection approach, that uses a novel Content-Defined Sampling method to generate a much smaller proxy hash set and then applies transforms on this small hash set. This reduces the calculations in the transform step from being the bottleneck. Meanwhile, Odess also leverages the faster Gear hash to generate rolling hashes. Thus, Odess greatly reduces the computational overhead for resemblance detection while achieving high detection accuracy and high compression ratio. Our evaluation results show that Odess is 5.4× (Finesse) and 26.9× (N-Transform) faster (on average) at generating features for resemblance detection. When considering an end-to-end data reduction storage system, Odess increases throughput by 1.36× (Finesse) and 2.76× (N-Transform) while maintaining the compression ratio of N-Transform and increasing the compression ratio 1.22× over Finesse.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 09ea4dc7-4b7a-4e32-8dbe-56122cc8eb23Cited by top-tier papers2
- Building a High-performance Fine-grained Deduplication Framework for Backup Storage with High Deduplication RatioXiangyu Zou, Wen Xia, Philip Shilane, Haijun Zhang et al.USENIX ATC 2022 · 36 citations
- LoopDelta: Embedding Locality-aware Opportunistic Delta Compression in Inline Deduplication for Highly Efficient Data ReductionYucheng Zhang, Hong Jiang, Dan Feng, Nan Jiang et al.USENIX ATC 2023 · 13 citations
Related papers
- Once Rolling Hashing is Enough: Exploiting Rolling Hash Reuse in Delta CompressionHaoliang Tan, Wenhao Ou, Xiangyu Zou, Cai Deng et al.EuroSys 2026 · 1 citation
- Palantir: Hierarchical Similarity Detection for Post-Deduplication Delta CompressionHongming Huang, Peng Wang, Qiang Su, Hong Xu et al.ASPLOS 2024 · 9 citations
- imDedup: A Lossless Deduplication Scheme to Eliminate Fine-grained Redundancy among ImagesCai Deng, Qi Chen, Xiangyu Zou, Erci Xu et al.ICDE 2022 · 16 citations
- DeepSketch: A New Machine Learning-Based Reference Search Technique for Post-Deduplication Delta CompressionJisung Park, Jeonggyun Kim, Yeseong Kim, Sungjin Lee et al.FAST 2022 · 40 citations
- The Dilemma between Deduplication and Locality: Can Both be Achieved?Xiangyu Zou, Jingsong Yuan, Philip Shilane, Wen Xia et al.FAST 2021 · 45 citations
