MESSI: In-Memory Data Series Indexing
Botao Peng, Panagiota Fatourou, Themis Palpanas
摘要
Data series similarity search is a core operation for several data series analysis applications across many different domains. However, the state-of-the-art techniques fail to deliver the time performance required for interactive exploration, or analysis of large data series collections. In this work, we propose MESSI, the first data series index designed for in-memory operation on modern hardware. Our index takes advantage of the modern hardware parallelization opportunities (i.e., SIMD instructions, multi-core and multi-socket architectures), in order to accelerate both index construction and similarity search processing times. Moreover, it benefits from a careful design in the setup and coordination of the parallel workers and data structures, so that it maximizes its performance for in-memory operations. Our experiments with synthetic and real datasets demonstrate that overall MESSI is up to 4× faster at index construction, and up to 11× faster at query answering than the state-of-the-art parallel approach. MESSI is the first to answer exact similarity search queries on 100GB datasets in 50msec (30-75msec across diverse datasets), which enables real-time, interactive data exploration on very large data series collections.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Elpis: Graph-Based Similarity Search for Scalable Data ScienceIlias Azizi, Karima Echihabi, Themis PalpanasVLDB 2023 · 被引用 67 次
- Data Series Progressive Similarity Search with Probabilistic Quality GuaranteesAnna Gogolou, Theophanis Tsandilas, Karima Echihabi, Anastasia Bezerianos 等SIGMOD 2020 · 被引用 38 次
- Graph-Based Vector Search: An Experimental Evaluation of the State-of-the-ArtIlias Azizi, Karima Echihabi, Themis PalpanasSIGMOD 2025 · 被引用 36 次
- DET-LSH: A Locality-Sensitive Hashing Scheme with Dynamic Encoding Tree for Approximate Nearest Neighbor SearchJiuqi Wei, Botao Peng, Xiaodong Lee, Themis PalpanasVLDB 2024 · 被引用 35 次
- Odyssey: A Journey in the Land of Distributed Data Series Similarity SearchManos Chatzakis, Panagiota Fatourou, Eleftherios Kosmas, Themis Palpanas 等VLDB 2023 · 被引用 26 次
它引用的顶会 Paper1
相关 Paper
- Hercules Against Data Series Similarity SearchKarima Echihabi, Panagiota Fatourou, Kostas Zoumpatianos, Themis Palpanas 等VLDB 2022 · 被引用 41 次
- Fast and Exact Similarity Search in Less than a Blink of an EyePatrick Schäfer, Jakob Brand, Ulf Leser, Botao Peng 等ICDE 2025 · 被引用 1 次
- Dumpy: A Compact and Adaptive Index for Large Data Series CollectionsZeyu Wang, Qitong Wang, Peng Wang, Themis Palpanas 等SIGMOD 2023 · 被引用 20 次
- DIDS: Double Indices and Double Summarizations for Fast Similarity SearchHan Hu, Jiye Qiu, Hongzhi Wang, Bin Liang 等VLDB 2024 · 被引用 2 次
- MS-Index: Fast Top-k Subsequence Search for Multivariate Time Series under Euclidean DistanceJens E. d'Hondt, Teun Kortekaas, Odysseas Papapetrou, Themis PalpanasVLDB 2026 · 被引用 1 次
