GAFormer: Enhancing Timeseries Transformers Through Group-Aware Embeddings
Jingyun Xiao, Ran Liu, Eva L. Dyer
Abstract
Analyzing multivariate time series is crucial in numerous domains, yet learning robust and generalizable representations within such datasets remains challenging due to complex inter-channel relationships and non-stationary dynamics. In this paper, we introduce a novel approach for learning data-adaptive position embeddings to incorporate learned spatial and temporal structure into transformer architectures. Our framework introduces group tokens and constructs an instancespecific group embedding (GE) layer that assigns input tokens to a select number of learned group tokens, thereby incorporating structural information into the learning process. Building on this, we propose a novel architecture, the Group-Aware Transformer (GAFormer), which integrates both spatial and temporal group embeddings to achieve state-of-the-art performance on various time series classification and regression tasks. Through evaluations on diverse time series datasets, we demonstrate that GE alone can significantly enhance the performance of several backbone models, and that the combination of spatial and temporal group embeddings allows GAFormer to surpass existing baselines. Moreover, our approach effectively discerns latent structures in data without prior knowledge of the spatial ordering of channels, leading to a more explainable decomposition of the spatial and temporal structure underlying complex timeseries datasets. Code is available at https://github.com/nerdslab/GAFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 205e022a-8a54-42d4-9c61-3fe093ea9bf3Cited by top-tier papers3
- Decentralized Attention Fails Centralized Signals: Rethinking Transformers for Medical Time SeriesGuoqi Yu, Juncheng Wang, Chen Yang, Jing Qin et al.ICLR 2026 · 6 citations
- Learning Fingerprints for Medical Time Series with Redundancy-Constrained Information MaximizationHuayu Li, ZhengXiao He, Xiwen Chen, Jingjing Wang et al.ICML 2026
- NeuroCLUS: A Foundation Model with Functional Clustering for Intracranial Neural DecodingHui Zheng, Haiteng WangICML 2026
Builds on13
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Rethinking and Improving Relative Position Encoding for Vision TransformerKan Wu, Houwen Peng, Minghao Chen, Jianlong Fu et al.ICCV 2021 · 427 citations
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang et al.ICLR 2023 · 406 citations
Related papers
- GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal TransformerShuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang et al.ICCV 2021 · 149 citations
- GaitCycFormer: Leveraging Gait Cycles and Transformers for Gait Emotion RecognitionQingyang Zeng, Lin ShangAAAI 2025 · 6 citations
- MedSpaformer: A Transferable Transformer with Multi-Granularity Token Sparsification for Medical Time Series ClassificationJiexia Ye, Weiqi Zhang, Ziyue Li, Jia Li et al.AAAI 2026 · 1 citation
- Crossformer: Transformer Utilizing Cross-Dimension Dependency for Multivariate Time Series ForecastingYunhao Zhang, Junchi YanICLR 2023
- CSformer: Combining Channel Independence and Mixing for Robust Multivariate Time Series ForecastingHaoxin Wang, Yipeng Mo, Kunlan Xiang, Nan Yin et al.AAAI 2025 · 9 citations
