Untangling tradeoffs between recurrence and self-attention in artificial neural networks
Giancarlo Kerg, Bhargav Kanuparthi, Anirudh Goyal, Kyle Goyette, Yoshua Bengio, Guillaume Lajoie
Abstract
Attention and self-attention mechanisms, are now central to state-of-the-art deep learning on sequential tasks. However, most recent progress hinges on heuristic approaches with limited understanding of attention's role in model optimization and computation, and rely on considerable memory and computational resources that scale poorly. In this work, we present a formal analysis of how self-attention affects gradient propagation in recurrent networks, and prove that it mitigates the problem of vanishing gradients when trying to capture long-term dependencies by establishing concrete bounds for gradient norms. Building on these results, we propose a relevancy screening mechanism, inspired by the cognitive process of memory consolidation, that allows for a scalable use of sparse self-attention with recurrence. While providing guarantees to avoid vanishing gradients, we use simple numerical experiments to demonstrate the tradeoffs in performance and computational resources by efficiently balancing attention and recurrence. Based on our results, we propose a concrete direction of research to improve scalability of attentive networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e9917ca-98de-4e4b-b262-4178fe86f174Cited by top-tier papers4
- Inductive Biases and Variable Creation in Self-Attention MechanismsBenjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Cyril ZhangICML 2022 · 154 citations
- Compositional Attention: Disentangling Search and RetrievalSarthak Mittal, Sharath Chandra Raparthy, Irina Rish, Yoshua Bengio et al.ICLR 2022 · 20 citations
- Abstractors and relational cross-attention: An inductive bias for explicit relational reasoning in TransformersAwni Altabaa, Taylor Whittington Webb, Jonathan D. Cohen, John LaffertyICLR 2024 · 13 citations
- Do Transformer World Models Give Better Policy Gradients?Michel Ma, Tianwei Ni, Clement Gehring, Pierluca D'Oro et al.ICML 2024 · 7 citations
Builds on1
Related papers
- On Vanishing Gradients, Over-Smoothing, and Over-Squashing in GNNs: Bridging Recurrent and Graph LearningAlvaro Arroyo, Alessio Gravina, Benjamin Gutteridge, Federico Barbero et al.NeurIPS 2025 · 58 citations
- Resolving the Timestep Scaling Paradox in Spiking Neural Networks with a Timestep-Scalable Neuron ModelBinghao Ye, Wenjuan Li, Dengfeng Xue, Bing Li et al.ICML 2026
- Continual learning in recurrent neural networksBenjamin Ehret, Christian Henning, Maria R. Cervera, Alexander Meulemans et al.ICLR 2021 · 4,433 citations
- Identifying nonlinear dynamical systems with multiple time scales and long-range dependenciesDominik Schmidt, Georgia Koppe, Zahra Monfared, Max Beutelspacher et al.ICLR 2021 · 41 citations
- Visual Attention Emerges from Recurrent Sparse ReconstructionBaifeng Shi, Yale Song, Neel Joshi, Trevor Darrell et al.ICML 2022 · 7 citations
