Representative Time Series Discovery for Data Exploration
Ge Lee, Shixun Huang, Zhifeng Bao, Yanchang Zhao
Abstract
In this work, we address the critical task of discovering representative time series in exploratory data mining. We define a representative time series, referred to as similarity-bounded representative time series, as one that represents other time series if their similarity meets a user-defined threshold. Building on this definition, we study the problem of finding the smallest set of such time series that can represent a specified proportion of all time series within the dataset. The representativeness of each similarity-bounded representative time series is controllable and determined by the specified level of similarity, and only the minimum number of such representatives needed to collectively represent the specified proportion of entire set are identified. Identifying representative time series over large-scale data in an efficient and effective manner facilitates exploratory data analysis and summary generation, serving a wide range of data exploration applications across diverse domains. We first prove the NP-hardness of this problem and propose a range of approximation methods with theoretical guarantees, and we refer to them as non-learning-based methods. While effective, these methods often excel in either running time or memory efficiency, but not both concurrently. To overcome these limitations, we further propose a learning-based method that simultaneously optimizes both time and memory efficiency. This method leverages novel data preparation and training strategies, providing adaptability to user-specified representativeness requirements with low memory usage and computational overhead. We conduct extensive experiments across four real-world datasets to demonstrate that our learning-based method is highly competitive with non-learning-based methods in terms of effectiveness (produces similar number of representative time series), while achieving significantly higher efficiency (up to 21× speedups) and lower memory consumption (saving up to 101× memory space).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a87918c-9333-442b-96a9-70fa2bd9443cBuilds on9
- SAND: Streaming Subsequence Anomaly DetectionPaul Boniol, John Paparrizos, Themis Palpanas, Michael J. FranklinVLDB 2021 · 128 citations
- Return of the Lernaean Hydra: Experimental Evaluation of Data Series Approximate Similarity SearchKarima Echihabi, Kostas Zoumpatianos, Themis Palpanas, Houda BenbrahimVLDB 2020 · 99 citations
- Data Diversification: A Simple Strategy For Neural Machine TranslationXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2020 · 75 citations
- KD-Box: Line-segment-based KD-tree for Interactive Exploration of Large-scale Time-Series DataYue Zhao, Yunhai Wang, Jian Zhang, Chi-Wing Fu et al.IEEE VIS 2021 · 45 citations
- Learning Representations for Incomplete Time Series ClusteringQianli Ma, Chuxin Chen, Sen Li, Garrison W. CottrellAAAI 2021 · 34 citations
Related papers
- Structure-Aware Abstraction of Hierarchical Time SeriesYihan Wu, Xuliang Zhu, Guozhong Li, Kai Wang et al.KDD 2026
- Efficient Learning-based Top-k Representative Similar Subtrajectory QueryKunming Wang, Shiyu Yang, Jiabao Jin, Peng Cheng et al.ICDE 2024 · 2 citations
- Representative Functional DependenciesQiongqiong Lin, Jingyan Sai, Jiazheng Song, Jinfei Liu et al.ICDE 2026
- SPARTAN: Data-Adaptive Symbolic Time-Series ApproximationFan Yang, John PaparrizosSIGMOD 2025 · 11 citations
- Multiscale Snapshots: Visual Analysis of Temporal Summaries in Dynamic GraphsEren Cakmak, Udo Schlegel, Dominik Jäckle, Daniel A. Keim et al.IEEE VIS 2020 · 20 citations
