Learning Hierarchical Polynomials of Multiple Nonlinear Features
Hengyu Fu, Zihao Wang, Eshaan Nichani, Jason D. Lee
摘要
In deep learning theory, a critical question is to understand how neural networks learn hierarchical features. In this work, we study the learning of hierarchical polynomials of multiple nonlinear features using three-layer neural networks. We examine a broad class of functions of the form f ⋆ = g ⋆ • p, where p : R d → R r represents multiple quadratic features with r ≪ d and g ⋆ : R r → R is a polynomial of degree p. This can be viewed as a nonlinear generalization of the multi-index model [Damian et al., 2022] , and also an expansion upon previous work that focused only on a single nonlinear feature, i.e. r = 1 [Nichani et al., 2023; Wang et al., 2023] . Our primary contribution shows that a three-layer neural network trained via layerwise gradient descent suffices for • complete recovery of the space spanned by the nonlinear features • efficient learning of the target function f ⋆ = g ⋆ • p or transfer learning of f = g • p with a different link function within O(d 4 ) samples and polynomial time. For such hierarchical targets, our result substantially improves the sample complexity Θ(d 2p ) of the kernel methods, demonstrating the power of efficient feature learning. It is important to highlight that our results leverage novel techniques and thus manage to go beyond all prior settings such as single-index and multi-index models as well as models depending just on one nonlinear feature, contributing to a more comprehensive understanding of feature learning in deep learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 被引用 128 次
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 · 被引用 119 次
相关 Paper
- Learning Hierarchical Polynomials with Three-Layer Neural NetworksZihao Wang, Eshaan Nichani, Jason D. LeeICLR 2024 · 被引用 7 次
- Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural NetworksEshaan Nichani, Alex Damian, Jason D. LeeNeurIPS 2023 · 被引用 24 次
- Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic LimitBohan Zhang, Zihao Wang, Hengyu Fu, Jason D. LeeICLR 2026 · 被引用 3 次
- Towards Understanding Hierarchical Learning: Benefits of Neural RepresentationsMinshuo Chen, Yu Bai, Jason D. Lee, Tuo Zhao 等NeurIPS 2020 · 被引用 61 次
- Neural network learns low-dimensional polynomials with SGD near the information-theoretic limitJason D. Lee, Kazusato Oko, Taiji Suzuki, Denny WuNeurIPS 2024 · 被引用 49 次
