End-to-end Algorithm Synthesis with Recurrent Networks: Extrapolation without Overthinking
Arpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam, Furong Huang, Micah Goldblum, Tom Goldstein
摘要
Machine learning systems perform well on pattern matching tasks, but their ability to perform algorithmic or logical reasoning is not well understood. One important reasoning capability is algorithmic extrapolation, in which models trained only on small/simple reasoning problems can synthesize complex strategies for large/complex problems at test time. Algorithmic extrapolation can be achieved through recurrent systems, which can be iterated many times to solve difficult reasoning problems. We observe that this approach fails to scale to highly complex problems because behavior degenerates when many iterations are applied -an issue we refer to as "overthinking." We propose a recall architecture that keeps an explicit copy of the problem instance in memory so that it cannot be forgotten. We also employ a progressive training routine that prevents the model from learning behaviors that are specific to iteration number and instead pushes it to learn behaviors that can be repeated indefinitely. These innovations prevent the overthinking problem, and enable recurrent systems to solve extremely hard extrapolation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer 等NeurIPS 2025 · 被引用 431 次
- Continuous Thought MachinesLuke Darlow, Ciaran Regan, Sebastian Risi, Jeffrey Seely 等NeurIPS 2025 · 被引用 22 次
- Computing a human-like reaction time metric from stable recurrent vision modelsLore Goetschalckx, Lakshmi Narasimhan Govindarajan, Alekh Karkada Ashok, Aarit Ahuja 等NeurIPS 2023 · 被引用 14 次
- Rethinking Deep Thinking: Stable Learning of Algorithms using Lipschitz ConstraintsJay Bear, Adam Prügel-Bennett, Jonathon S. HareNeurIPS 2024 · 被引用 10 次
它引用的顶会 Paper4
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent NetworksAvi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang 等NeurIPS 2021 · 被引用 133 次
- The Uncanny Similarity of Recurrence and DepthAvi Schwarzschild, Arjun Gupta, Amin Ghiasi, Micah Goldblum 等ICLR 2022 · 被引用 11 次
- Differentiable Adaptive Computation Time for Visual ReasoningCristóbal Eyzaguirre, Álvaro SotoCVPR 2020
相关 Paper
- On Logical Extrapolation for Mazes with Recurrent and Implicit NetworksBrandon Knutson, Amandin Chyba Rabeendran, Michael I. Ivanitskiy, Jordan Pettyjohn 等AAAI 2026 · 被引用 7 次
- Recognizing and Verifying Mathematical Equations using Multiplicative Differential Neural UnitsAnkur Mali, Alexander G. Ororbia II, Daniel Kifer, C. Lee GilesAAAI 2021 · 被引用 16 次
- InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language ModelsYuchen Yan, Yongliang Shen, Yang Liu, Jin Jiang 等ICLR 2026 · 被引用 48 次
- Learning Iterative Reasoning through Energy MinimizationYilun Du, Shuang Li, Joshua B. Tenenbaum, Igor MordatchICML 2022 · 被引用 37 次
- RLAD: Training LLMs to Discover Abstractions for Solving Reasoning ProblemsYuxiao Qu, Anikait Singh, Yoonho Lee, Amrith Setlur 等ICLR 2026 · 被引用 25 次
