Odess: Speeding up Resemblance Detection for Redundancy Elimination by Fast Content-Defined Sampling
Xiangyu Zou, Cai Deng, Wen Xia, Philip Shilane, Haoliang Tan, Haijun Zhang, Xuan Wang
摘要
Multiple data reduction techniques have been investigated to lower storage costs for a wide variety of customers. In this work, we focus on similarity-based delta compression, which calculates and stores the difference of very similar, but non-duplicate, chunks in storage systems. Delta compression is often implemented along with deduplication and has been shown to achieve a much higher compression ratio. Currently, the N-Transform method is the most popular and widely-used approach to generate features for data content (e.g. chunks) to detect similar candidates (and then apply delta compression). For delta compression systems, though, the throughput of N-Transform is often the bottleneck. Finesse is a high throughput variant of N-Transform, but it suffers from lower detection accuracy and compression ratio. The computation overhead of N-Transform consists of two parts: calculating the rolling hash across data and applying time-consuming transforms on each hash. In this work, we propose Odess, a fast resemblance detection approach, that uses a novel Content-Defined Sampling method to generate a much smaller proxy hash set and then applies transforms on this small hash set. This reduces the calculations in the transform step from being the bottleneck. Meanwhile, Odess also leverages the faster Gear hash to generate rolling hashes. Thus, Odess greatly reduces the computational overhead for resemblance detection while achieving high detection accuracy and high compression ratio. Our evaluation results show that Odess is 5.4× (Finesse) and 26.9× (N-Transform) faster (on average) at generating features for resemblance detection. When considering an end-to-end data reduction storage system, Odess increases throughput by 1.36× (Finesse) and 2.76× (N-Transform) while maintaining the compression ratio of N-Transform and increasing the compression ratio 1.22× over Finesse.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Building a High-performance Fine-grained Deduplication Framework for Backup Storage with High Deduplication RatioXiangyu Zou, Wen Xia, Philip Shilane, Haijun Zhang 等USENIX ATC 2022 · 被引用 36 次
- LoopDelta: Embedding Locality-aware Opportunistic Delta Compression in Inline Deduplication for Highly Efficient Data ReductionYucheng Zhang, Hong Jiang, Dan Feng, Nan Jiang 等USENIX ATC 2023 · 被引用 13 次
相关 Paper
- Once Rolling Hashing is Enough: Exploiting Rolling Hash Reuse in Delta CompressionHaoliang Tan, Wenhao Ou, Xiangyu Zou, Cai Deng 等EuroSys 2026 · 被引用 1 次
- Palantir: Hierarchical Similarity Detection for Post-Deduplication Delta CompressionHongming Huang, Peng Wang, Qiang Su, Hong Xu 等ASPLOS 2024 · 被引用 9 次
- imDedup: A Lossless Deduplication Scheme to Eliminate Fine-grained Redundancy among ImagesCai Deng, Qi Chen, Xiangyu Zou, Erci Xu 等ICDE 2022 · 被引用 16 次
- DeepSketch: A New Machine Learning-Based Reference Search Technique for Post-Deduplication Delta CompressionJisung Park, Jeonggyun Kim, Yeseong Kim, Sungjin Lee 等FAST 2022 · 被引用 40 次
- The Dilemma between Deduplication and Locality: Can Both be Achieved?Xiangyu Zou, Jingsong Yuan, Philip Shilane, Wen Xia 等FAST 2021 · 被引用 45 次
