The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization
Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber
Abstract
Despite progress across a broad range of applications, Transformers have limited success in systematic generalization. The situation is especially frustrating in the case of algorithmic tasks, where they often fail to find intuitive solutions that route relevant information to the right node/operation at the right time in the grid represented by Transformer columns. To facilitate the learning of useful control flow, we propose two modifications to the Transformer architecture, copy gate and geometric attention. Our novel Neural Data Router (NDR) achieves 100% length generalization accuracy on the classic compositional table lookup task, as well as near-perfect accuracy on the simple arithmetic task and a new variant of ListOps testing for generalization across computational depths. NDR's attention and gating patterns tend to be interpretable as an intuitive form of neural routing. Our code is public. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c5e05a1-192f-44b8-9386-d14eafb191d9Cited by top-tier papers24
- Block-Recurrent TransformersDeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer et al.NeurIPS 2022 · 163 citations
- Going Beyond Linear Transformers with Recurrent Fast Weight ProgrammersKazuki Irie, Imanol Schlag, Róbert Csordás, Jürgen SchmidhuberNeurIPS 2021 · 101 citations
- MoEUT: Mixture-of-Experts Universal TransformersRóbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts et al.NeurIPS 2024 · 65 citations
- SALSA: Attacking Lattice Cryptography with TransformersEmily Wenger, Mingjie Chen, François Charton, Kristin E. LauterNeurIPS 2022 · 61 citations
- How Far Can Transformers Reason? The Globality Barrier and Inductive ScratchpadEmmanuel Abbe, Samy Bengio, Aryo Lotfi, Colin Sandon et al.NeurIPS 2024 · 52 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
Related papers
- Neural Execution Engines: Learning to Execute SubroutinesYujun Yan, Kevin Swersky, Danai Koutra, Parthasarathy Ranganathan et al.NeurIPS 2020 · 47 citations
- Looped Transformers for Length GeneralizationYing Fan, Yilun Du, Kannan Ramchandran, Kangwook LeeICLR 2025
- Extrapolation by Association: Length Generalization Transfer In TransformersZiyang Cai, Nayoung Lee, Avi Schwarzschild, Samet Oymak et al.NeurIPS 2025 · 13 citations
- Mastering Symbolic Operations: Augmenting Language Models with Compiled Neural NetworksYixuan Weng, Minjun Zhu, Fei Xia, Bin Li et al.ICLR 2024 · 14 citations
- INViT: A Generalizable Routing Problem Solver with Invariant Nested View TransformerHan Fang, Zhihao Song, Paul Weng, Yutong BanICML 2024 · 39 citations
