Representative Time Series Discovery for Data Exploration
Ge Lee, Shixun Huang, Zhifeng Bao, Yanchang Zhao
摘要
In this work, we address the critical task of discovering representative time series in exploratory data mining. We define a representative time series, referred to as similarity-bounded representative time series, as one that represents other time series if their similarity meets a user-defined threshold. Building on this definition, we study the problem of finding the smallest set of such time series that can represent a specified proportion of all time series within the dataset. The representativeness of each similarity-bounded representative time series is controllable and determined by the specified level of similarity, and only the minimum number of such representatives needed to collectively represent the specified proportion of entire set are identified. Identifying representative time series over large-scale data in an efficient and effective manner facilitates exploratory data analysis and summary generation, serving a wide range of data exploration applications across diverse domains. We first prove the NP-hardness of this problem and propose a range of approximation methods with theoretical guarantees, and we refer to them as non-learning-based methods. While effective, these methods often excel in either running time or memory efficiency, but not both concurrently. To overcome these limitations, we further propose a learning-based method that simultaneously optimizes both time and memory efficiency. This method leverages novel data preparation and training strategies, providing adaptability to user-specified representativeness requirements with low memory usage and computational overhead. We conduct extensive experiments across four real-world datasets to demonstrate that our learning-based method is highly competitive with non-learning-based methods in terms of effectiveness (produces similar number of representative time series), while achieving significantly higher efficiency (up to 21× speedups) and lower memory consumption (saving up to 101× memory space).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- SAND: Streaming Subsequence Anomaly DetectionPaul Boniol, John Paparrizos, Themis Palpanas, Michael J. FranklinVLDB 2021 · 被引用 128 次
- Return of the Lernaean Hydra: Experimental Evaluation of Data Series Approximate Similarity SearchKarima Echihabi, Kostas Zoumpatianos, Themis Palpanas, Houda BenbrahimVLDB 2020 · 被引用 99 次
- Data Diversification: A Simple Strategy For Neural Machine TranslationXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2020 · 被引用 75 次
- KD-Box: Line-segment-based KD-tree for Interactive Exploration of Large-scale Time-Series DataYue Zhao, Yunhai Wang, Jian Zhang, Chi-Wing Fu 等IEEE VIS 2021 · 被引用 45 次
- Learning Representations for Incomplete Time Series ClusteringQianli Ma, Chuxin Chen, Sen Li, Garrison W. CottrellAAAI 2021 · 被引用 34 次
相关 Paper
- Structure-Aware Abstraction of Hierarchical Time SeriesYihan Wu, Xuliang Zhu, Guozhong Li, Kai Wang 等KDD 2026
- Efficient Learning-based Top-k Representative Similar Subtrajectory QueryKunming Wang, Shiyu Yang, Jiabao Jin, Peng Cheng 等ICDE 2024 · 被引用 2 次
- Representative Functional DependenciesQiongqiong Lin, Jingyan Sai, Jiazheng Song, Jinfei Liu 等ICDE 2026
- SPARTAN: Data-Adaptive Symbolic Time-Series ApproximationFan Yang, John PaparrizosSIGMOD 2025 · 被引用 11 次
- Multiscale Snapshots: Visual Analysis of Temporal Summaries in Dynamic GraphsEren Cakmak, Udo Schlegel, Dominik Jäckle, Daniel A. Keim 等IEEE VIS 2020 · 被引用 20 次
