Time-Series Clustering: A Comprehensive Study of Data Mining, Machine Learning, and Deep Learning Methods
John Paparrizos, Bogireddy Sai Prasanna Teja
摘要
Time-series clustering is a key task in time series analysis, enabling unsupervised data exploration and often serving as a subroutine for other tasks. Despite decades of active cross-disciplinary research, benchmarking of time-series clustering methods has received limited attention. Existing studies have (i) excluded popular methods and entire method classes; (ii) used a narrow range of distance measures; (iii) evaluated only a few datasets; (iv) lacked statistical validation; (v) had poor reproducibility; or (vi) relied on questionable evaluation setups. The rise of deep learning—especially foundation models claiming broad generalization—further emphasizes the need for comprehensive evaluation, as their role in time-series clustering remains largely untested. To address these gaps, we evaluate 84 time-series clustering methods across 10 method classes from data mining, machine learning, and deep learning. Our analysis spans 128 time-series datasets and uses rigorous statistical methods. Within a fair comparison framework, we (i) identify the top-performing method in each class; (ii) highlight previously overlooked, high-performing classes; (iii) challenge assumptions about elastic distance measures; (iv) refute the claimed superiority of deep learning methods, including foundation models; (v) expose reproducibility issues; (vi) analyze performance variation across dataset properties; and (vii) assess scalability. Our findings reveal an illusion of progress: no method significantly outperforms the decade-old k -Shape method. Still, we highlight a deep learning-based approach with notable promise. Our results provide a strong benchmark for advancing time-series clustering, and we have open-sourced our work to support future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Beyond Compression: A Comprehensive Evaluation of Lossless Floating-Point CompressionKaisei Hishida, Chunwei Liu, John Paparrizos, Aaron J. ElmoreVLDB 2025 · 被引用 8 次
- The Power of Anomaly Detection in Predictive Maintenance: [Experiments & Analysis]Anastasios Papadopoulos, Apostolos Giannoulidis, Anastasios Gounaris, John PaparrizosSIGMOD 2026 · 被引用 4 次
它引用的顶会 Paper20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun 等NeurIPS 2023 · 被引用 1,178 次
- Structural Deep Clustering NetworkDeyu Bo, Xiao Wang, Chuan Shi, Meiqi Zhu 等WWW 2020 · 被引用 645 次
- MOMENT: A Family of Open Time-series Foundation ModelsMononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai 等ICML 2024 · 被引用 442 次
- Volume Under the Surface: A New Accuracy Evaluation Measure for Time-Series Anomaly DetectionJohn Paparrizos, Paul Boniol, Themis Palpanas, Ruey S. Tsay 等VLDB 2022 · 被引用 171 次
相关 Paper
- MUFASA: Fast and Accurate Multivariate Time-Series ClusteringHaojun Li, John PaparrizosSIGMOD 2026 · 被引用 2 次
- TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting MethodsXiangfei Qiu, Jilin Hu, Lekui Zhou, Xingjian Wu 等VLDB 2024 · 被引用 292 次
- TAB: Unified Benchmarking of Time Series Anomaly Detection MethodsXiangfei Qiu, Zhe Li, Wanghui Qiu, Shiyan Hu 等VLDB 2025 · 被引用 57 次
- Debunking Four Long-Standing Misconceptions of Time-Series Distance MeasuresJohn Paparrizos, Chunwei Liu, Aaron J. Elmore, Michael J. FranklinSIGMOD 2020 · 被引用 56 次
- Efficient Deep Embedded Subspace ClusteringJinyu Cai, Jicong Fan, Wenzhong Guo, Shiping Wang 等CVPR 2022 · 被引用 127 次
