Wavy Transformer
Satoshi Noguchi, Yoshinobu Kawahara
Abstract
Transformers have achieved remarkable success across natural language processing (NLP) and computer vision (CV). However, deep transformer models often suffer from an over-smoothing issue, in which token representations converge to similar values as they pass through successive transformer blocks. In this paper, we establish an equivalence between the hidden-state dynamics induced by stacked attention layers and graph neural diffusion on a complete graph. From this perspective, over-smoothing can be interpreted as a consequence of the dissipative nature of the underlying diffusion dynamics. Motivated by this physical interpretation, we propose Wavy Transformer, which consists of a novel attention layer based on second-order wavy dynamics. We also introduce a feedforward network and a normalization layer designed to preserve the physical state-velocity relationship under the chain rule, thereby extending the transformer architecture. We further validate our proposed techniques on various transformer models for NLP, CV, and sparse-graph tasks. The results consistently demonstrate that Wavy Transformer improves performance with minimal additional parameters and no extra hyperparameter tuning. Source code and models are available at https://github.com/noguchisatoshi/Wavy-Transformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
Related papers
- Mitigating Over-smoothing in Transformers via Regularized Nonlocal FunctionalsTam Nguyen, Tan M. Nguyen, Richard G. BaraniukNeurIPS 2023 · 46 citations
- Revisiting Over-smoothing in BERT from the Perspective of GraphHan Shi, Jiahui Gao, Hang Xu, Xiaodan Liang et al.ICLR 2022 · 92 citations
- Graph Convolutions Enrich the Self-Attention in Transformers!Jeongwhan Choi, Hyowon Wi, Jayoung Kim, Yehjin Shin et al.NeurIPS 2024 · 24 citations
- Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence TransformersYukun Zhang, Xueqing ZhouEMNLP 2025
- Demystifying Oversmoothing in Attention-Based Graph Neural NetworksXinyi Wu, Amir Ajorlou, Zihui Wu, Ali JadbabaieNeurIPS 2023 · 86 citations
