Towards Understanding Hierarchical Learning: Benefits of Neural Representations
Minshuo Chen, Yu Bai, Jason D. Lee, Tuo Zhao, Huan Wang, Caiming Xiong, Richard Socher
Abstract
Deep neural networks can empirically perform efficient hierarchical learning, in which the layers learn useful representations of the data. However, how they make use of the intermediate representations are not explained by recent theories that relate them to "shallow learners" such as kernels. In this work, we demonstrate that intermediate neural representations add more flexibility to neural networks and can be advantageous over raw inputs. We consider a fixed, randomly initialized neural network as a representation function fed into another trainable network. When the trainable network is the quadratic Taylor model of a wide two-layer network, we show that neural representation can achieve improved sample complexities compared with the raw input: For learning a low-rank degree- polynomial () in dimension, neural representation requires only samples, while the best-known sample complexity upper bound for the raw input is . We contrast our result with a lower bound showing that neural representations do not improve over the raw input (in the infinite width limit), when the trainable network is instead a neural tangent kernel. Our results characterize when neural representations are beneficial, and may provide a new perspective on why depth is important in deep learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3605fb1a-318f-4ef8-aadf-8e2992a947e7Cited by top-tier papers26
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 · 119 citations
- The staircase property: How hierarchical structure can guide deep learningEmmanuel Abbe, Enric Boix-Adserà, Matthew S. Brennan, Guy Bresler et al.NeurIPS 2021 · 74 citations
- Feature learning in deep classifiers through Intermediate Neural CollapseAkshay Rangamani, Marius Lindegaard, Tomer Galanti, Tomaso A. PoggioICML 2023 · 64 citations
- A Theoretical Analysis on Feature Learning in Neural Networks: Emergence from Inputs and Advantage over Fixed FeaturesZhenmei Shi, Junyi Wei, Yingyu LiangICLR 2022 · 58 citations
- Local Signal Adaptivity: Provable Feature Learning in Neural Networks Beyond KernelsStefani Karp, Ezra Winston, Yuanzhi Li, Aarti SinghNeurIPS 2021 · 38 citations
Builds on4
- Generalization Error Bounds of Gradient Descent for Learning Over-Parameterized Deep ReLU NetworksYuan Cao, Quanquan GuAAAI 2020 · 168 citations
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 128 citations
- Asymptotics of Wide Networks from Feynman DiagramsEthan Dyer, Guy Gur-AriICLR 2020 · 127 citations
Related papers
- Learning Hierarchical Polynomials with Three-Layer Neural NetworksZihao Wang, Eshaan Nichani, Jason D. LeeICLR 2024 · 7 citations
- Learning Hierarchical Polynomials of Multiple Nonlinear FeaturesHengyu Fu, Zihao Wang, Eshaan Nichani, Jason D. LeeICLR 2025
- Sharp Representation Theorems for ReLU Networks with Precise Dependence on DepthGuy Bresler, Dheeraj NagarajNeurIPS 2020 · 27 citations
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 62 citations
- Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic LimitBohan Zhang, Zihao Wang, Hengyu Fu, Jason D. LeeICLR 2026 · 3 citations
