The Uncanny Similarity of Recurrence and Depth
Avi Schwarzschild, Arjun Gupta, Amin Ghiasi, Micah Goldblum, Tom Goldstein
摘要
It is widely believed that deep neural networks contain layer specialization, wherein neural networks extract hierarchical features representing edges and patterns in shallow layers and complete objects in deeper layers. Unlike common feed-forward models that have distinct filters at each layer, recurrent networks reuse the same parameters at various depths. In this work, we observe that recurrent models exhibit the same hierarchical behaviors and the same performance benefits with depth as feed-forward networks despite reusing the same filters at every recurrence. By training models of various feed-forward and recurrent architectures on several datasets for image classification as well as maze solving, we show that recurrent networks have the ability to closely emulate the behavior of non-recurrent deep models, often doing so with far fewer parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- End-to-end Algorithm Synthesis with Recurrent Networks: Extrapolation without OverthinkingArpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam 等NeurIPS 2022 · 被引用 54 次
- CoBERL: Contrastive BERT for Reinforcement LearningAndrea Banino, Adrià Puigdomènech Badia, Jacob C. Walker, Tim Scholtes 等ICLR 2022 · 被引用 41 次
- Block Recurrent Dynamics in Vision TransformersMozes Jacobs, Thomas Fel, Richard Hakim, Alessandra Brondetta 等ICLR 2026 · 被引用 17 次
- Short-Term Plasticity Neurons Learning to Learn and ForgetHector Garcia Rodriguez, Qinghai Guo, Timoleon MoraitisICML 2022 · 被引用 15 次
- Adaptive recurrent vision performs zero-shot computation scaling to unseen difficulty levelsVijay Veerabadran, Srinivas Ravishankar, Yuan Tang, Ritik Raina 等NeurIPS 2023 · 被引用 11 次
它引用的顶会 Paper3
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 被引用 613 次
- Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversionHongxu Yin, Pavlo Molchanov, José M. Álvarez, Zhizhong Li 等CVPR 2020
相关 Paper
- Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent NetworksAvi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang 等NeurIPS 2021 · 被引用 133 次
- The Master Key Filters Hypothesis: Deep Filters Are GeneralZahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu GrosuAAAI 2025 · 被引用 3 次
- Machines of Finite Depth: Towards a Formalization of Neural NetworksPietro Vertechi, Mattia G. BergomiAAAI 2023 · 被引用 2 次
- Multi-Task Recurrent Modular NetworksDongkuan Xu, Wei Cheng, Xin Dong, Bo Zong 等AAAI 2021 · 被引用 2 次
- Path Independent Equilibrium Models Can Better Exploit Test-Time ComputationCem Anil, Ashwini Pokle, Kaiqu Liang, Johannes Treutlein 等NeurIPS 2022 · 被引用 43 次
