Decentralized Attention Fails Centralized Signals: Rethinking Transformers for Medical Time Series
Guoqi Yu, Juncheng Wang, Chen Yang, Jing Qin, Angelica I Aviles-Rivero, Shujun Wang
摘要
Accurate analysis of Medical time series (MedTS) data, such as Electroencephalography (EEG) and Electrocardiography (ECG), plays a pivotal role in healthcare applications, including the diagnosis of brain and heart diseases. MedTS data typically exhibits two critical patterns: temporal dependencies within individual channels and channel dependencies across multiple channels. While recent advances in deep learning have leveraged Transformer-based models to effectively capture temporal dependencies, they often struggle to model channel dependencies. This limitation stems from a structural mismatch: MedTS signals are inherently centralized, whereas the Transformer's attention is decentralized, making it less effective at capturing global synchronization and unified waveform patterns. To bridge this gap, we propose CoTAR (Core Token Aggregation-Redistribution), a centralized MLP-based module tailored to replace the decentralized attention. Instead of allowing all tokens to interact directly, as in attention, CoTAR introduces a global core token that acts as a proxy to facilitate the inter-token interaction, thereby enforcing a centralized aggregation and redistribution strategy. This design not only better aligns with the centralized nature of MedTS signals but also reduces computational complexity from quadratic to linear. Experiments on five benchmarks validate the superiority of our method in both effectiveness and efficiency, achieving up to a 11.6% improvement on the APAVA dataset, with merely 33% memory usage and 20% inference time compared to the previous state-of-the-art. Code and all training scripts are available in this Link.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 5,824 次
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang 等ICML 2022 · 被引用 2,912 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
相关 Paper
- Medformer: A Multi-Granularity Patching Transformer for Medical Time-Series ClassificationYihe Wang, Nan Huang, Taida Li, Yujun Yan 等NeurIPS 2024 · 被引用 158 次
- MedSpaformer: A Transferable Transformer with Multi-Granularity Token Sparsification for Medical Time Series ClassificationJiexia Ye, Weiqi Zhang, Ziyue Li, Jia Li 等AAAI 2026 · 被引用 1 次
- Reading Between the Channels: Knowledge-Augmented Medical Time Series ClassificationXiaoyan Yuan, Wei Wang, Junxin Chen, Xiping HuACM MM 2025 · 被引用 1 次
- TimeMIL: Advancing Multivariate Time Series Classification via a Time-aware Multiple Instance LearningXiwen Chen, Peijie Qiu, Wenhui Zhu, Huayu Li 等ICML 2024 · 被引用 31 次
- A Closer Look at Transformers for Time Series Forecasting: Understanding Why They Work and Where They StruggleYu Chen, Nathalia Céspedes, Payam M. BarnaghiICML 2025
