Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel Perspective
Kaixuan Huang, Yuqing Wang, Molei Tao, Tuo Zhao
摘要
Deep residual networks (ResNets) have demonstrated better generalization performance than deep feedforward networks (FFNets). However, the theory behind such a phenomenon is still largely unknown. This paper studies this fundamental problem in deep learning from a so-called "neural tangent kernel" perspective. Specifically, we first show that under proper conditions, as the width goes to infinity, training deep ResNets can be viewed as learning reproducing kernel functions with some kernel function. We then compare the kernel of deep ResNets with that of deep FFNets and discover that the class of functions induced by the kernel of FFNets is asymptotically not learnable, as the depth goes to infinity. In contrast, the class of functions induced by the kernel of ResNets does not exhibit such degeneracy. Our discovery partially justifies the advantages of deep ResNets over deep FFNets in generalization abilities. Numerical results are provided to support our claim.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Gradient Starvation: A Learning Proclivity in Neural NetworksMohammad Pezeshki, Sékou-Oumar Kaba, Yoshua Bengio, Aaron C. Courville 等NeurIPS 2021 · 被引用 378 次
- The staircase property: How hierarchical structure can guide deep learningEmmanuel Abbe, Enric Boix-Adserà, Matthew S. Brennan, Guy Bresler 等NeurIPS 2021 · 被引用 74 次
- What can a Single Attention Layer Learn? A Study Through the Random Features LensHengyu Fu, Tianyu Guo, Yu Bai, Song MeiNeurIPS 2023 · 被引用 47 次
- Neural Tangent Kernel Beyond the Infinite-Width Limit: Effects of Depth and InitializationMariia Seleznova, Gitta KutyniokICML 2022 · 被引用 34 次
- A Neural Tangent Kernel Perspective of GANsJean-Yves Franceschi, Emmanuel de Bézenac, Ibrahim Ayed, Mickaël Chen 等ICML 2022 · 被引用 29 次
相关 Paper
- Branch Scaling Manifests as Implicit Architectural Regularization for Improving Generalization in Overparameterized ResNetsZixiong Yu, Guhan Chen, Jianfa Lai, Bohan Li 等ICML 2026 · 被引用 2 次
- On the Random Conjugate Kernel and Neural Tangent KernelZhengmian Hu, Heng HuangICML 2021 · 被引用 15 次
- The Recurrent Neural Tangent KernelSina Alemohammad, Zichao Wang, Randall Balestriero, Richard G. BaraniukICLR 2021 · 被引用 6 次
- Disentangling Trainability and Generalization in Deep Neural NetworksLechao Xiao, Jeffrey Pennington, Samuel Stern SchoenholzICML 2020 · 被引用 91 次
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 被引用 62 次
