Global-Lens Transformers: Adaptive Token Mixing for Dynamic Link Prediction
Tao Zou, Chengfeng Wu, Tianxi Liao, Junchen Ye, Bowen Du
Abstract
Dynamic graph learning plays a pivotal role in modeling evolving relationships over time, especially for temporal link prediction tasks in domains such as traffic systems, social networks, and recommendation platforms. While Transformer-based models have demonstrated strong performance by capturing long-range temporal dependencies, their reliance on self-attention results in quadratic complexity with respect to sequence length, limiting scalability on high-frequency or large-scale graphs. In this work, we revisit the necessity of self-attention in dynamic graph modeling. Inspired by recent findings that attribute the success of Transformers more to their architectural design than attention itself, we propose GLFormer, a novel attention-free Transformer-style framework for dynamic graphs. GLFormer introduces an adaptive token mixer that performs context-aware local aggregation based on interaction order and time intervals. To capture long-term dependencies, we further design a hierarchical aggregation module that expands the temporal receptive field by stacking local token mixers across layers. Experiments on six widely used dynamic graph benchmarks show that GLFormer achieves competitive or superior performance, which reveals that attention-free architectures can match or surpass Transformer baselines in dynamic graph settings with significantly improved efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Pyraformer: Low-Complexity Pyramidal Attention for Long-Range Time Series Modeling and ForecastingShizhan Liu, Hang Yu, Cong Liao, Jianguo Li et al.ICLR 2022 · 975 citations
- Inductive representation learning on temporal graphsDa Xu, Chuanwei Ruan, Evren Körpeoglu, Sushant Kumar et al.ICLR 2020 · 901 citations
- Rethinking Graph Transformers with Spectral AttentionDevin Kreuzer, Dominique Beaini, William L. Hamilton, Vincent Létourneau et al.NeurIPS 2021 · 854 citations
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 653 citations
Related papers
- Towards Better Dynamic Graph Learning: New Architecture and Unified LibraryLe Yu, Leilei Sun, Bowen Du, Weifeng LvNeurIPS 2023 · 323 citations
- TFWaveFormer: Temporal-Frequency Collaborative Multi-level Wavelet Transformer for Dynamic Link PredictionHantong Feng, Yonggang Wu, Duxin Chen, Wenwu YuWWW 2026
- TIDFormer: Exploiting Temporal and Interactive Dynamics Makes A Great Dynamic Graph TransformerJie Peng, Zhewei Wei, Yuhang YeKDD 2025 · 2 citations
- On the Feasibility of Simple Transformer for Dynamic Graph ModelingYuxia Wu, Yuan Fang, Lizi LiaoWWW 2024 · 49 citations
- LLGformer: Learnable Long-range Graph Transformer for Traffic Flow PredictionDi Jin, Cuiying Huo, Jiayi Shi, Dongxiao He et al.WWW 2025 · 14 citations
