Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical Systems
Subhabrata Dutta, Tanya Gautam, Soumen Chakrabarti, Tanmoy Chakraborty
Abstract
The Transformer and its variants have been proven to be efficient sequence learners in many different domains. Despite their staggering success, a critical issue has been the enormous number of parameters that must be trained (ranging from 10 7 to 10 11 ) along with the quadratic complexity of dot-product attention. In this work, we investigate the problem of approximating the two central components of the Transformer -multi-head self-attention and point-wise feed-forward transformation, with reduced parameter space and computational complexity. We build upon recent developments in analyzing deep neural networks as numerical solvers of ordinary differential equations. Taking advantage of an analogy between Transformer stages and the evolution of a dynamical system of multiple interacting particles, we formulate a temporal evolution scheme, TransEvolve, to bypass costly dot-product attention over multiple stacked layers. We perform exhaustive experiments with TransEvolve on well-known encoder-decoder as well as encoder-only tasks. We observe that the degree of approximation (or inversely, the degree of parameter reduction) has different effects on the performance, depending on the task. While in the encoder-decoder regime, TransEvolve delivers performances comparable to the original Transformer, in encoder-only tasks it consistently outperforms Transformer along with several subsequent variants. Code is available in: https://github.com/LCS2-IIITD/TransEvolve .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2164cd8-e4d5-4a87-ac83-671db02ff91eCited by top-tier papers15
- LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut ModulationAhmadreza Jeddi, Marco Ciccone, Babak TaatiICLR 2026 · 54 citations
- Seeing the forest and the tree: Building representations of both individual and collective dynamics with transformersRan Liu, Mehdi Azabou, Max Dabagia, Jingyun Xiao et al.NeurIPS 2022 · 28 citations
- Perceptrons and Localization of Attention’s Mean-Field LandscapeAntonio Álvarez López, Borjan Geshkovski, Domènec Ruiz-BaletICML 2026 · 7 citations
- Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient LearningBei Li, Tong Zheng, Rui Wang, Jiahao Liu et al.NeurIPS 2024 · 5 citations
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale ProblemsSaeed Amizadeh, Sara Abdali, Yinheng Li, Kazuhito KoishidaNeurIPS 2025 · 2 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
Related papers
- Understanding the Expressive Power and Mechanisms of Transformer for Sequence ModelingMingze Wang, Weinan ENeurIPS 2024 · 32 citations
- On the approximation properties of recurrent encoder-decoder architecturesZhong Li, Haotian Jiang, Qianxiao LiICLR 2022 · 8 citations
- The Effect of Attention Head Count on Transformer ApproximationPenghao Yu, Haotian Jiang, Zeyu Bao, Ruoxi Yu et al.ICLR 2026 · 5 citations
- Transolver Is a Linear Transformer: Revisiting Physics-Attention Through the Lens of Linear AttentionWenjie Hu, Sidun Liu, Peng Qiao, Zhenglun Sun et al.AAAI 2026 · 2 citations
- Recurrent Attention for Neural Machine TranslationJiali Zeng, Shuangzhi Wu, Yongjing Yin, Yufan Jiang et al.EMNLP 2021
