MUFASA: Fast and Accurate Multivariate Time-Series Clustering
Haojun Li, John Paparrizos
Abstract
Time-series clustering is a core analytical technique for data exploration and as a subroutine in downstream tasks. Although deep learning methods have surged recently, studies show that traditional algorithms, such as k -Shape, still achieve state-of-the-art performance, revealing an illusion of progress. k -Shape alternates between (i) cluster assignment using the Shape-based distance (SBD) and (ii) centroid update while preserving scale and temporal alignment. However, its cubic-time centroid computation limits scalability to long sequences. Despite advances in univariate time-series (UTS) clustering, multivariate (MTS) clustering remains underexplored—multiple channels complicate inter-channel dependency modeling and amplify the accuracy–runtime trade-off. To overcome these challenges, we introduce FASA and its multivariate extension MUFASA: scalable, accurate, and parameter-light methods. FASA derives a closed-form solution that minimizes within-cluster SBD distance in linear time, while MUFASA (i) introduces SBD-D, a distance measure identifying a global temporal alignment across channels, and (ii) introduces a centroid update that jointly estimates all channels to capture temporal and inter-channel dependencies efficiently. To demonstrate the effectiveness, we conduct the most comprehensive evaluation to date in the MTS clustering area—covering seven MTS distances and 21 clustering algorithms on 30 UEA MTS datasets—and benchmark FASA on 128 UCR UTS datasets against eight baselines. Results show that (i) SBD-D matches the accuracy of elastic distances while being on average two to four orders of magnitude faster, and (ii) MUFASA outperforms all scalable traditional, deep learning, and foundation models in accuracy while matching non-scalable ones at much lower runtime. Notably, FASA matches k -Shape's accuracy while exhibiting better scalability, especially with increasing length. Overall, MUFASA and FASA deliver efficient and accurate solutions for MTS and UTS clustering, respectively, paving the way for future progress in time-series analytics.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 52581500-2013-481d-bde4-04087a3c8cafCited by top-tier papers2
- The Power of Anomaly Detection in Predictive Maintenance: [Experiments & Analysis]Anastasios Papadopoulos, Apostolos Giannoulidis, Anastasios Gounaris, John PaparrizosSIGMOD 2026 · 4 citations
- HYDRA: A Multi-Level Hierarchy-Driven Approach for Robust Anomaly Detection in Time SeriesMingyi Huang, Qinghua Liu, Paul Boniol, John PaparrizosSIGMOD 2026 · 4 citations
Related papers
- Time-Series Clustering: A Comprehensive Study of Data Mining, Machine Learning, and Deep Learning MethodsJohn Paparrizos, Bogireddy Sai Prasanna TejaVLDB 2025 · 13 citations
- In-Database Time Series ClusteringYunxiang Su, Kenny Ye Liang, Shaoxu SongSIGMOD 2025 · 3 citations
- SVP-T: A Shape-Level Variable-Position Transformer for Multivariate Time Series ClassificationRundong Zuo, Guozhong Li, Byron Choi, Sourav S. Bhowmick et al.AAAI 2023 · 53 citations
- ShapeNet: A Shapelet-Neural Network Approach for Multivariate Time Series ClassificationGuozhong Li, Byron Choi, Jianliang Xu, Sourav S. Bhowmick et al.AAAI 2021 · 177 citations
- Time2Feat: Learning Interpretable Representations for Multivariate Time Series ClusteringAngela Bonifati, Francesco Del Buono, Francesco Guerra, Donato TianoVLDB 2023 · 27 citations
