DIDS: Double Indices and Double Summarizations for Fast Similarity Search
Han Hu, Jiye Qiu, Hongzhi Wang, Bin Liang, Songling Zou
Abstract
Data series has been one of the significant data forms in various applications. It becomes imperative to devise a data series index that supports both approximate and exact similarity searches for large data series collections in high-dimensional metric spaces. The state-of-the-art works employ summarizations and indices to reduce the accesses to the data series. However, we discover two significant flaws that severely limit performance enhancement. Firstly, the state-of-the-art works often employ segment-based summarizations, whose lower bound distances decrease significantly when representing a data series collection, resulting in numerous invalid accesses. Secondly, the disk-based indices for the exact search mainly rely on tree-based indices, which results in low-quality approximate answers, consequently impacting the exact search. To address these problems, we propose a novel solution, Double Indices and Double Summarizations (DIDS). Besides segment-based summarizations, DIDS introduces reference-point-based summarizations to improve the pruning rate by the sorted-based representation strategy. Moreover, DIDS employs reference points and a cost model to cluster similar data series, and uses a graph-based approach to interconnect various regions, enhancing approximate search capabilities. We conduct experiments on extensive datasets, validating the superior search performance of DIDS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- SPANN: Highly-efficient Billion-scale Approximate Nearest Neighborhood SearchQi Chen, Bing Zhao, Haidong Wang, Mingqin Li et al.NeurIPS 2021 · 219 citations
- Return of the Lernaean Hydra: Experimental Evaluation of Data Series Approximate Similarity SearchKarima Echihabi, Kostas Zoumpatianos, Themis Palpanas, Houda BenbrahimVLDB 2020 · 99 citations
- HVS: Hierarchical Graph Structure Based on Voronoi Diagrams for Solving Approximate Nearest Neighbor SearchKejing Lu, Mineichi Kudo, Chuan Xiao, Yoshiharu IshikawaVLDB 2022 · 70 citations
- Elpis: Graph-Based Similarity Search for Scalable Data ScienceIlias Azizi, Karima Echihabi, Themis PalpanasVLDB 2023 · 67 citations
- Hercules Against Data Series Similarity SearchKarima Echihabi, Panagiota Fatourou, Kostas Zoumpatianos, Themis Palpanas et al.VLDB 2022 · 41 citations
Related papers
- LeaFi: Data Series Indexes on Steroids with Learned FiltersQitong Wang, Ioana Ileana, Themis PalpanasSIGMOD 2025 · 9 citations
- Constructing Compact Time Series Index for Efficient Window Query ProcessingJing Zhao, Peng Wang, Bo Tang, Lu Liu et al.ICDE 2022 · 1 citation
- MESSI: In-Memory Data Series IndexingBotao Peng, Panagiota Fatourou, Themis PalpanasICDE 2020 · 38 citations
- CLIMBER: Pivot-Based Approximate Similarity Search Over Big Data SeriesLiang Zhang, Mohamed Y. Eltabakh, Elke A. Rundensteiner, Khalid AlnuaimICDE 2024 · 1 citation
- Dumpy: A Compact and Adaptive Index for Large Data Series CollectionsZeyu Wang, Qitong Wang, Peng Wang, Themis Palpanas et al.SIGMOD 2023 · 20 citations
