Adaptive recurrent vision performs zero-shot computation scaling to unseen difficulty levels
Vijay Veerabadran, Srinivas Ravishankar, Yuan Tang, Ritik Raina, Virginia de Sa
摘要
Humans solving algorithmic (or) reasoning problems typically exhibit solution times that grow as a function of problem difficulty. Adaptive recurrent neural networks have been shown to exhibit this property for various language-processing tasks. However, little work has been performed to assess whether such adaptive computation can also enable vision models to extrapolate solutions beyond their training distribution's difficulty level, with prior work focusing on very simple tasks. In this study, we investigate a critical functional role of such adaptive processing using recurrent neural networks: to dynamically scale computational resources conditional on input requirements that allow for zero-shot generalization to novel difficulty levels not seen during training using two challenging visual reasoning tasks: PathFinder and Mazes. We combine convolutional recurrent neural networks (ConvRNNs) with a learnable halting mechanism based on (Graves, 2016). We explore various implementations of such adaptive ConvRNNs (AdRNNs) ranging from tying weights across layers to more sophisticated biologically inspired recurrent networks that possess lateral connections and gating. We show that 1) AdRNNs learn to dynamically halt processing early (or late) to solve easier (or harder) problems, 2) these RNNs zero-shot generalize to more difficult problem settings not shown during training by dynamically increasing the number of recurrent iterations at test time. Our study provides modeling evidence supporting the hypothesis that recurrent processing enables the functional advantage of adaptively allocating compute resources conditional on input requirements and hence allowing generalization to harder difficulty levels of a visual reasoning problem without training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Flexible Context-Driven Sensory Processing in Dynamical Vision ModelsLakshmi Narasimhan Govindarajan, Abhiram Iyer, Valmiki Kothare, Ila FieteNeurIPS 2024 · 被引用 1 次
- NeuralSolver: Learning Algorithms For Consistent and Efficient Extrapolation Across General TasksBernardo Esteves, Miguel Vasco, Francisco S. MeloNeurIPS 2024 · 被引用 1 次
- Understanding and Improving Length Generalization in Recurrent ModelsRicardo Buitrago Ruiz, Albert GuICML 2025
它引用的顶会 Paper6
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent NetworksAvi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang 等NeurIPS 2021 · 被引用 133 次
- Stable and expressive recurrent vision modelsDrew Linsley, Alekh Karkada Ashok, Lakshmi Narasimhan Govindarajan, Rex G. Liu 等NeurIPS 2020 · 被引用 56 次
- Improving Anytime Prediction with Parallel Cascaded Networks and a Temporal-Difference LossMichael L. Iuzzolino, Michael C. Mozer, Samy BengioNeurIPS 2021 · 被引用 15 次
- The Uncanny Similarity of Recurrence and DepthAvi Schwarzschild, Arjun Gupta, Amin Ghiasi, Micah Goldblum 等ICLR 2022 · 被引用 11 次
相关 Paper
- On Logical Extrapolation for Mazes with Recurrent and Implicit NetworksBrandon Knutson, Amandin Chyba Rabeendran, Michael I. Ivanitskiy, Jordan Pettyjohn 等AAAI 2026 · 被引用 7 次
- Learning Iterative Reasoning through Energy MinimizationYilun Du, Shuang Li, Joshua B. Tenenbaum, Igor MordatchICML 2022 · 被引用 37 次
- Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual ReasoningFan Shi, Bin Li, Xiangyang XueICML 2025
- End-to-end Algorithm Synthesis with Recurrent Networks: Extrapolation without OverthinkingArpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam 等NeurIPS 2022 · 被引用 54 次
- Equilibrium Reasoners: Learning Attractors Enables Scalable ReasoningBenhao Huang, Zhengyang Geng, Zico KolterICML 2026 · 被引用 5 次
