MTM: A Multi-Scale Token Mixing Transformer for Irregular Multivariate Time Series Classification
Shuhan Zhong, Weipeng Zhuo, Sizhe Song, Guanyao Li, Zhongyi Yu, S.-H. Gary Chan
Abstract
Irregular multivariate time series (IMTS) is characterized by the lack of synchronized observations across its different channels. In this paper, we point out that this channel-wise asynchrony can lead to poor channel-wise modeling of existing deep learning methods. To overcome this limitation, we propose MTM, a multi-scale token mixing transformer for the classification of IMTS. We find that the channel-wise asynchrony can be alleviated by down-sampling the time series to coarser timescales, and propose to incorporate a masked concat pooling in MTM that gradually down-samples IMTS to enhance the channel-wise attention modules. Meanwhile, we propose a novel channel-wise token mixing mechanism which proactively chooses important tokens from one channel and mixes them with other channels, to further boost the channel-wise learning of our model. Through extensive experiments on real-world datasets and comparison with state-of-the-art methods, we demonstrate that MTM consistently achieves the best performance on all the benchmarks, with improvements of up to 3.8% in AUPRC for classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5be99758-eb49-43f4-9940-a54fc4ff2378Builds on21
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
- TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting MethodsXiangfei Qiu, Jilin Hu, Lekui Zhou, Xingjian Wu et al.VLDB 2024 · 292 citations
- HiTANet: Hierarchical Time-Aware Attention Networks for Risk Prediction on Electronic Health RecordsJunyu Luo, Muchao Ye, Cao Xiao, Fenglong MaKDD 2020 · 187 citations
Related papers
- TimeCHEAT: A Channel Harmony Strategy for Irregularly Sampled Multivariate Time Series AnalysisJiexi Liu, Meng Cao, Songcan ChenAAAI 2025 · 16 citations
- A Multi-Scale Decomposition MLP-Mixer for Time Series AnalysisShuhan Zhong, Sizhe Song, Weipeng Zhuo, Guanyao Li et al.VLDB 2024 · 48 citations
- QuITE: Query-Based Irregular Time Series EmbeddingJunghoon LimICML 2026
- Learning Recursive Multi-Scale Representations for Irregular Multivariate Time Series ForecastingBoyuan Li, Zhen Liu, Yicheng Luo, Qianli MaICLR 2026
- TimeMIL: Advancing Multivariate Time Series Classification via a Time-aware Multiple Instance LearningXiwen Chen, Peijie Qiu, Wenhui Zhu, Huayu Li et al.ICML 2024 · 31 citations
