The Uncanny Similarity of Recurrence and Depth
Avi Schwarzschild, Arjun Gupta, Amin Ghiasi, Micah Goldblum, Tom Goldstein
Abstract
It is widely believed that deep neural networks contain layer specialization, wherein neural networks extract hierarchical features representing edges and patterns in shallow layers and complete objects in deeper layers. Unlike common feed-forward models that have distinct filters at each layer, recurrent networks reuse the same parameters at various depths. In this work, we observe that recurrent models exhibit the same hierarchical behaviors and the same performance benefits with depth as feed-forward networks despite reusing the same filters at every recurrence. By training models of various feed-forward and recurrent architectures on several datasets for image classification as well as maze solving, we show that recurrent networks have the ability to closely emulate the behavior of non-recurrent deep models, often doing so with far fewer parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13f8fdbe-dd6d-45d8-a913-a27f46917c20Cited by top-tier papers5
- End-to-end Algorithm Synthesis with Recurrent Networks: Extrapolation without OverthinkingArpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam et al.NeurIPS 2022 · 54 citations
- CoBERL: Contrastive BERT for Reinforcement LearningAndrea Banino, Adrià Puigdomènech Badia, Jacob C. Walker, Tim Scholtes et al.ICLR 2022 · 41 citations
- Block Recurrent Dynamics in Vision TransformersMozes Jacobs, Thomas Fel, Richard Hakim, Alessandra Brondetta et al.ICLR 2026 · 17 citations
- Short-Term Plasticity Neurons Learning to Learn and ForgetHector Garcia Rodriguez, Qinghai Guo, Timoleon MoraitisICML 2022 · 15 citations
- Adaptive recurrent vision performs zero-shot computation scaling to unseen difficulty levelsVijay Veerabadran, Srinivas Ravishankar, Yuan Tang, Ritik Raina et al.NeurIPS 2023 · 11 citations
Builds on3
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversionHongxu Yin, Pavlo Molchanov, José M. Álvarez, Zhizhong Li et al.CVPR 2020
Related papers
- Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent NetworksAvi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang et al.NeurIPS 2021 · 133 citations
- The Master Key Filters Hypothesis: Deep Filters Are GeneralZahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu GrosuAAAI 2025 · 3 citations
- Machines of Finite Depth: Towards a Formalization of Neural NetworksPietro Vertechi, Mattia G. BergomiAAAI 2023 · 2 citations
- Multi-Task Recurrent Modular NetworksDongkuan Xu, Wei Cheng, Xin Dong, Bo Zong et al.AAAI 2021 · 2 citations
- Path Independent Equilibrium Models Can Better Exploit Test-Time ComputationCem Anil, Ashwini Pokle, Kaiqu Liang, Johannes Treutlein et al.NeurIPS 2022 · 43 citations
