Physical vs. Logical Indexing with IDEA: Inverted Deduplication-Aware Index
Asaf Levi, Philip Shilane, Sarai Sheinvald, Gala Yadgar
摘要
In the realm of information retrieval, the need to maintain reliable term-indexing has grown more acute in recent years, with vast amounts of ever-growing online data used for data mining and natural language processing, and searched by a large number of search-engine users. At the same time, an increasing portion of primary storage systems employ data deduplication, where duplicate logical data chunks are replaced with references to a unique physical copy. We show that indexing deduplicated data with deduplication-oblivious mechanisms might result in extreme inefficiencies: the index size would increase in proportion to the logical data size, regardless of its duplication ratio, consuming excessive storage and memory and slowing down lookups. In addition, the logically sequential accesses during index creation would be transformed into random and redundant accesses to the physical chunks. Indeed, to the best of our knowledge, term indexing is not supported by any deduplicating storage system. In this article, we propose the design of a deduplication-aware term-index that addresses these challenges. IDEA maps terms to the unique chunks that contain them, and maps each chunk to the files in which it is contained. This basic design concept improves the index performance and can support advanced functionalities such as inline indexing, result ranking, and proximity search. Our prototype implementation based on Lucene (the search engine at the core of Elasticsearch) shows that IDEA can reduce the index size and indexing time by up to 73% and 94%, respectively, and reduce term-lookup latency by up to 82% and 59% for single and multi-term queries, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- Pre-training via ParaphrasingMike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan 等NeurIPS 2020 · 被引用 165 次
- GoSeed: Generating an Optimal Seeding Plan for Deduplicated StorageAviv Nachman, Gala Yadgar, Sarai SheinvaldFAST 2020 · 被引用 19 次
- The what, The from, and The to: The Migration Games in Deduplicated SystemsRoei Kisous, Ariel Kolikant, Abhinav Duggal, Sarai Sheinvald 等FAST 2022 · 被引用 12 次
- DedupSearch: Two-Phase Deduplication Aware Keyword SearchNadav Elias, Philip Shilane, Sarai Sheinvald, Gala YadgarFAST 2022 · 被引用 5 次
相关 Paper
- FinerDedup: Sifting Fingerprints for Efficient Data Deduplication on Mobile DevicesXianzhang Chen, Xingjie Zhou, Wei Li, Xi Yu 等DAC 2024 · 被引用 2 次
- The Dilemma between Deduplication and Locality: Can Both be Achieved?Xiangyu Zou, Jingsong Yuan, Philip Shilane, Wen Xia 等FAST 2021 · 被引用 45 次
- Austere Flash Caching with Deduplication and CompressionQiuping Wang, Jinhong Li, Wen Xia, Erik Kruus 等USENIX ATC 2020 · 被引用 26 次
- IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model InferenceWeijian Chen, Shuibing He, Haoyang Qu, Ruidong Zhang 等FAST 2025 · 被引用 40 次
- Don't Maintain Twice, It's Alright: Merged Metadata Management in Deduplication File System with GogetaFSYanqi Pan, Wen Xia, Erci Xu, Hao Huang 等FAST 2025 · 被引用 6 次
