Deep Learning Embeddings for Data Series Similarity Search
Qitong Wang, Themis Palpanas
Abstract
A key operation for the (increasingly large) data series collection analysis is similarity search. According to recent studies, SAX-based indexes offer state-of-the-art performance for similarity search tasks. However, their performance lags under high-frequency, weakly correlated, excessively noisy, or other dataset-specific properties. In this work, we propose Deep Embedding Approximation (DEA), a novel family of data series summarization techniques based on deep neural networks. Moreover, we describe SEAnet, a novel architecture especially designed for learning DEA, that introduces the Sum of Squares preservation property into the deep network design. Finally, we propose a new sampling strategy, SEASam, that allows SEAnet to effectively train on massive datasets. Comprehensive experiments on 7 diverse synthetic and real datasets verify the advantages of DEA learned using SEAnet, when compared to other state-of-the-art traditional and DEA solutions, in providing high-quality data series summarizations and similarity search results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b1f6e1c-3837-456f-8ae0-d19cd06e2172Cited by top-tier papers15
- SAND: Streaming Subsequence Anomaly DetectionPaul Boniol, John Paparrizos, Themis Palpanas, Michael J. FranklinVLDB 2021 · 128 citations
- Elpis: Graph-Based Similarity Search for Scalable Data ScienceIlias Azizi, Karima Echihabi, Themis PalpanasVLDB 2023 · 67 citations
- Hercules Against Data Series Similarity SearchKarima Echihabi, Panagiota Fatourou, Kostas Zoumpatianos, Themis Palpanas et al.VLDB 2022 · 41 citations
- Graph-Based Vector Search: An Experimental Evaluation of the State-of-the-ArtIlias Azizi, Karima Echihabi, Themis PalpanasSIGMOD 2025 · 36 citations
- dCAM: Dimension-wise Class Activation Map for Explaining Multivariate Data Series ClassificationPaul Boniol, Mohammed Meftah, Emmanuel Remy, Themis PalpanasSIGMOD 2022 · 27 citations
Builds on3
- SAND: Streaming Subsequence Anomaly DetectionPaul Boniol, John Paparrizos, Themis Palpanas, Michael J. FranklinVLDB 2021 · 128 citations
- Return of the Lernaean Hydra: Experimental Evaluation of Data Series Approximate Similarity SearchKarima Echihabi, Kostas Zoumpatianos, Themis Palpanas, Houda BenbrahimVLDB 2020 · 99 citations
- Series2Graph: Graph-based Subsequence Anomaly Detection for Time SeriesPaul Boniol, Themis PalpanasVLDB 2020
Related papers
- Fast and Exact Similarity Search in Less than a Blink of an EyePatrick Schäfer, Jakob Brand, Ulf Leser, Botao Peng et al.ICDE 2025 · 1 citation
- leSAX Index: A Learned SAX Representation Index for Time Series Similarity SearchGuozhong Li, Byron Choi, Rundong Zuo, Sourav S. Bhowmick et al.ICDE 2025 · 3 citations
- DIDS: Double Indices and Double Summarizations for Fast Similarity SearchHan Hu, Jiye Qiu, Hongzhi Wang, Bin Liang et al.VLDB 2024 · 2 citations
- Learned Cardinality Estimation for Similarity QueriesJi Sun, Guoliang Li, Nan TangSIGMOD 2021 · 40 citations
- MESSI: In-Memory Data Series IndexingBotao Peng, Panagiota Fatourou, Themis PalpanasICDE 2020 · 38 citations
