Learning Representations for Incomplete Time Series Clustering
Qianli Ma, Chuxin Chen, Sen Li, Garrison W. Cottrell
Abstract
Time-series clustering is an essential unsupervised technique for data analysis, applied to many real-world fields, such as medical analysis and DNA microarray. Existing clustering methods are usually based on the assumption that the data is complete. However, time series in real-world applications often contain missing values. Traditional strategy (imputing first and then clustering) does not optimize the imputation and clustering process as a whole, which not only makes per- formance dependent on the combination of imputation and clustering methods but also fails to achieve satisfactory re- sults. How to best improve the clustering performance on incomplete time series remains a challenge. This paper pro- poses a novel unsupervised temporal representation learning model, named Clustering Representation Learning on Incom- plete time-series data (CRLI). CRLI jointly optimizes the im- putation and clustering process to impute more discrimina- tive values for clustering and make the learned representa- tions possessed good clustering property. Also, to reduce the error propagation from imputation to clustering, we introduce a discriminator to make the distribution of imputation values close to the true one and train CRLI in an alternating train- ing manner. An experiment conducted on eight real-world in- complete time-series datasets shows that CRLI outperforms existing methods. We demonstrates the effectiveness of the learned representations and the convergence of the model through visualization analysis. Moreover, we reveal that the joint training strategy can impute values close to the true ones in those important sub-sequences, and impute more discrim- inative values in those less important sub-sequences at the same time, making the imputed sequence cluster-friendly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 131a8aa3-6d0e-4ffc-839a-d4b65d2702beCited by top-tier papers2
- Representative Time Series Discovery for Data ExplorationGe Lee, Shixun Huang, Zhifeng Bao, Yanchang ZhaoVLDB 2025 · 2 citations
- Rethinking Time-Series Imputation as Conditional Inference along Temporal EvolutionYu Fan, Yang Yang, guo yufan, Huazhong Yang et al.ICML 2026
Related papers
- Self-Representation Subspace Clustering for Incomplete Multi-view DataJiyuan Liu, Xinwang Liu, Yi Zhang, Pei Zhang et al.ACM MM 2021 · 98 citations
- Integrating Sequence and Image Modeling in Irregular Medical Time Series Through Self-Supervised LearningLiuqing Chen, Shuhong Xiao, Shixian Ding, Shanhai Hu et al.AAAI 2025 · 3 citations
- Attribute-Missing Graph Clustering NetworkWenxuan Tu, Renxiang Guan, Sihang Zhou, Chuan Ma et al.AAAI 2024 · 51 citations
- COMPLETER: Incomplete Multi-View Clustering via Contrastive PredictionYijie Lin, Yuanbiao Gou, Zitao Liu, Boyun Li et al.CVPR 2021
- Generative Semi-supervised Learning for Multivariate Time Series ImputationXiaoye Miao, Yangyang Wu, Jun Wang, Yunjun Gao et al.AAAI 2021 · 212 citations
