Dumpy: A Compact and Adaptive Index for Large Data Series Collections
Zeyu Wang, Qitong Wang, Peng Wang, Themis Palpanas, Wei Wang
Abstract
Data series indexes are necessary for managing and analyzing the increasing amounts of data series collections that are nowadays available. These indexes support both exact and approximate similarity search, with approximate search providing high-quality results within milliseconds, which makes it very attractive for certain modern applications. Reducing the pre-processing (i.e., index building) time and improving the accuracy of search results are two major challenges. DSTree and the iSAX index family are state-of-the-art solutions for this problem. However, DSTree suffers from long index building times, while iSAX suffers from low search accuracy. In this paper, we identify two problems of the iSAX index family that adversely affect the overall performance. First, we observe the presence of a proximity-compactness trade-off related to the index structure design (i.e., the node fanout degree), significantly limiting the efficiency and accuracy of the resulting index. Second, a skewed data distribution will negatively affect the performance of iSAX. To overcome these problems, we propose Dumpy, an index that employs a novel multi-ary data structure with an adaptive node splitting algorithm and an efficient building workflow. Furthermore, we devise Dumpy-Fuzzy as a variant of Dumpy which further improves search accuracy by proper duplication of series. Experiments with a variety of large, real datasets demonstrate that the Dumpy solutions achieve considerably better efficiency, scalability and search accuracy than its competitors. This paper was published in SIGMOD'23.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fb2d13a-c0f2-40f3-ba0d-59c22a39225fCited by top-tier papers18
- RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor SearchJianyang Gao, Cheng LongSIGMOD 2024 · 83 citations
- Elpis: Graph-Based Similarity Search for Scalable Data ScienceIlias Azizi, Karima Echihabi, Themis PalpanasVLDB 2023 · 67 citations
- Graph-Based Vector Search: An Experimental Evaluation of the State-of-the-ArtIlias Azizi, Karima Echihabi, Themis PalpanasSIGMOD 2025 · 36 citations
- DET-LSH: A Locality-Sensitive Hashing Scheme with Dynamic Encoding Tree for Approximate Nearest Neighbor SearchJiuqi Wei, Botao Peng, Xiaodong Lee, Themis PalpanasVLDB 2024 · 35 citations
- Odyssey: A Journey in the Land of Distributed Data Series Similarity SearchManos Chatzakis, Panagiota Fatourou, Eleftherios Kosmas, Themis Palpanas et al.VLDB 2023 · 26 citations
Builds on11
- A Comprehensive Survey and Experimental Comparison of Graph-Based Approximate Nearest Neighbor SearchMengzhao Wang, Xiaoliang Xu, Qiang Yue, Yuxiang WangVLDB 2021 · 354 citations
- SPANN: Highly-efficient Billion-scale Approximate Nearest Neighborhood SearchQi Chen, Bing Zhao, Haidong Wang, Mingqin Li et al.NeurIPS 2021 · 219 citations
- Return of the Lernaean Hydra: Experimental Evaluation of Data Series Approximate Similarity SearchKarima Echihabi, Kostas Zoumpatianos, Themis Palpanas, Houda BenbrahimVLDB 2020 · 99 citations
- Elpis: Graph-Based Similarity Search for Scalable Data ScienceIlias Azizi, Karima Echihabi, Themis PalpanasVLDB 2023 · 67 citations
- MESSI: In-Memory Data Series IndexingBotao Peng, Panagiota Fatourou, Themis PalpanasICDE 2020 · 38 citations
Related papers
- CLIMBER: Pivot-Based Approximate Similarity Search Over Big Data SeriesLiang Zhang, Mohamed Y. Eltabakh, Elke A. Rundensteiner, Khalid AlnuaimICDE 2024 · 1 citation
- DIDS: Double Indices and Double Summarizations for Fast Similarity SearchHan Hu, Jiye Qiu, Hongzhi Wang, Bin Liang et al.VLDB 2024 · 2 citations
- LeaFi: Data Series Indexes on Steroids with Learned FiltersQitong Wang, Ioana Ileana, Themis PalpanasSIGMOD 2025 · 9 citations
- Hercules Against Data Series Similarity SearchKarima Echihabi, Panagiota Fatourou, Kostas Zoumpatianos, Themis Palpanas et al.VLDB 2022 · 41 citations
- leSAX Index: A Learned SAX Representation Index for Time Series Similarity SearchGuozhong Li, Byron Choi, Rundong Zuo, Sourav S. Bhowmick et al.ICDE 2025 · 3 citations
