Learning Hierarchical Polynomials of Multiple Nonlinear Features
Hengyu Fu, Zihao Wang, Eshaan Nichani, Jason D. Lee
Abstract
In deep learning theory, a critical question is to understand how neural networks learn hierarchical features. In this work, we study the learning of hierarchical polynomials of multiple nonlinear features using three-layer neural networks. We examine a broad class of functions of the form f ⋆ = g ⋆ • p, where p : R d → R r represents multiple quadratic features with r ≪ d and g ⋆ : R r → R is a polynomial of degree p. This can be viewed as a nonlinear generalization of the multi-index model [Damian et al., 2022] , and also an expansion upon previous work that focused only on a single nonlinear feature, i.e. r = 1 [Nichani et al., 2023; Wang et al., 2023] . Our primary contribution shows that a three-layer neural network trained via layerwise gradient descent suffices for • complete recovery of the space spanned by the nonlinear features • efficient learning of the target function f ⋆ = g ⋆ • p or transfer learning of f = g • p with a different link function within O(d 4 ) samples and polynomial time. For such hierarchical targets, our result substantially improves the sample complexity Θ(d 2p ) of the kernel methods, demonstrating the power of efficient feature learning. It is important to highlight that our results leverage novel techniques and thus manage to go beyond all prior settings such as single-index and multi-index models as well as models depending just on one nonlinear feature, contributing to a more comprehensive understanding of feature learning in deep learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb146924-4c14-4db7-b455-0dad2ca19a35Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam et al.NeurIPS 2020 · 245 citations
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 · 173 citations
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 128 citations
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 · 119 citations
Related papers
- Learning Hierarchical Polynomials with Three-Layer Neural NetworksZihao Wang, Eshaan Nichani, Jason D. LeeICLR 2024 · 7 citations
- Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural NetworksEshaan Nichani, Alex Damian, Jason D. LeeNeurIPS 2023 · 24 citations
- Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic LimitBohan Zhang, Zihao Wang, Hengyu Fu, Jason D. LeeICLR 2026 · 3 citations
- Towards Understanding Hierarchical Learning: Benefits of Neural RepresentationsMinshuo Chen, Yu Bai, Jason D. Lee, Tuo Zhao et al.NeurIPS 2020 · 61 citations
- Neural network learns low-dimensional polynomials with SGD near the information-theoretic limitJason D. Lee, Kazusato Oko, Taiji Suzuki, Denny WuNeurIPS 2024 · 49 citations
