Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent Networks
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum, Tom Goldstein
摘要
Deep neural networks are powerful machines for visual pattern recognition, but reasoning tasks that are easy for humans may still be difficult for neural models. Humans possess the ability to extrapolate reasoning strategies learned on simple problems to solve harder examples, often by thinking for longer. For example, a person who has learned to solve small mazes can easily extend the very same search techniques to solve much larger mazes by spending more time. In computers, this behavior is often achieved through the use of algorithms, which scale to arbitrarily hard problem instances at the cost of more computation. In contrast, the sequential computing budget of feed-forward neural networks is limited by their depth, and networks trained on simple problems have no way of extending their reasoning to accommodate harder problems. In this work, we show that recurrent networks trained to solve simple problems with few recurrent steps can indeed solve much more complex problems simply by performing additional recurrences during inference. We demonstrate this algorithmic behavior of recurrent networks on prefix sum computation, mazes, and chess. In all three domains, networks trained on simple problem instances are able to extend their reasoning abilities at test time simply by "thinking for longer."
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 被引用 898 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 等ICML 2024 · 被引用 443 次
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei 等ICLR 2023 · 被引用 318 次
它引用的顶会 Paper3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Aligning Superhuman AI with Human Behavior: Chess as a Model SystemReid McIlroy-Young, Siddhartha Sen, Jon M. Kleinberg, Ashton AndersonKDD 2020 · 被引用 77 次
- Differentiable Adaptive Computation Time for Visual ReasoningCristóbal Eyzaguirre, Álvaro SotoCVPR 2020
相关 Paper
- End-to-end Algorithm Synthesis with Recurrent Networks: Extrapolation without OverthinkingArpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam 等NeurIPS 2022 · 被引用 54 次
- Learning Iterative Reasoning through Energy MinimizationYilun Du, Shuang Li, Joshua B. Tenenbaum, Igor MordatchICML 2022 · 被引用 37 次
- Adaptive recurrent vision performs zero-shot computation scaling to unseen difficulty levelsVijay Veerabadran, Srinivas Ravishankar, Yuan Tang, Ritik Raina 等NeurIPS 2023 · 被引用 11 次
- On Logical Extrapolation for Mazes with Recurrent and Implicit NetworksBrandon Knutson, Amandin Chyba Rabeendran, Michael I. Ivanitskiy, Jordan Pettyjohn 等AAAI 2026 · 被引用 7 次
- Recognizing and Verifying Mathematical Equations using Multiplicative Differential Neural UnitsAnkur Mali, Alexander G. Ororbia II, Daniel Kifer, C. Lee GilesAAAI 2021 · 被引用 16 次
