End-to-end Algorithm Synthesis with Recurrent Networks: Extrapolation without Overthinking
Arpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam, Furong Huang, Micah Goldblum, Tom Goldstein
Abstract
Machine learning systems perform well on pattern matching tasks, but their ability to perform algorithmic or logical reasoning is not well understood. One important reasoning capability is algorithmic extrapolation, in which models trained only on small/simple reasoning problems can synthesize complex strategies for large/complex problems at test time. Algorithmic extrapolation can be achieved through recurrent systems, which can be iterated many times to solve difficult reasoning problems. We observe that this approach fails to scale to highly complex problems because behavior degenerates when many iterations are applied -an issue we refer to as "overthinking." We propose a recall architecture that keeps an explicit copy of the problem instance in memory so that it cannot be forgotten. We also employ a progressive training routine that prevents the model from learning behaviors that are specific to iteration number and instead pushes it to learn behaviors that can be repeated indefinitely. These innovations prevent the overthinking problem, and enable recurrent systems to solve extremely hard extrapolation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8e6c022-42bf-4564-ad0c-536ddf398abfCited by top-tier papers25
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer et al.NeurIPS 2025 · 431 citations
- Continuous Thought MachinesLuke Darlow, Ciaran Regan, Sebastian Risi, Jeffrey Seely et al.NeurIPS 2025 · 22 citations
- Computing a human-like reaction time metric from stable recurrent vision modelsLore Goetschalckx, Lakshmi Narasimhan Govindarajan, Alekh Karkada Ashok, Aarit Ahuja et al.NeurIPS 2023 · 14 citations
- Rethinking Deep Thinking: Stable Learning of Algorithms using Lipschitz ConstraintsJay Bear, Adam Prügel-Bennett, Jonathon S. HareNeurIPS 2024 · 10 citations
Builds on4
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent NetworksAvi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang et al.NeurIPS 2021 · 133 citations
- The Uncanny Similarity of Recurrence and DepthAvi Schwarzschild, Arjun Gupta, Amin Ghiasi, Micah Goldblum et al.ICLR 2022 · 11 citations
- Differentiable Adaptive Computation Time for Visual ReasoningCristóbal Eyzaguirre, Álvaro SotoCVPR 2020
Related papers
- On Logical Extrapolation for Mazes with Recurrent and Implicit NetworksBrandon Knutson, Amandin Chyba Rabeendran, Michael I. Ivanitskiy, Jordan Pettyjohn et al.AAAI 2026 · 7 citations
- Recognizing and Verifying Mathematical Equations using Multiplicative Differential Neural UnitsAnkur Mali, Alexander G. Ororbia II, Daniel Kifer, C. Lee GilesAAAI 2021 · 16 citations
- InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language ModelsYuchen Yan, Yongliang Shen, Yang Liu, Jin Jiang et al.ICLR 2026 · 48 citations
- Learning Iterative Reasoning through Energy MinimizationYilun Du, Shuang Li, Joshua B. Tenenbaum, Igor MordatchICML 2022 · 37 citations
- RLAD: Training LLMs to Discover Abstractions for Solving Reasoning ProblemsYuxiao Qu, Anikait Singh, Yoonho Lee, Amrith Setlur et al.ICLR 2026 · 25 citations
