When Expressivity Meets Trainability: Fewer than Neurons Can Work
Jiawei Zhang, Yushun Zhang, Mingyi Hong, Ruoyu Sun, Zhi-Quan Luo
摘要
Modern neural networks are often quite wide, causing large memory and computation costs. It is thus of great interest to train a narrower network. However, training narrow neural nets remains a challenging task. We ask two theoretical questions: Can narrow networks have as strong expressivity as wide ones? If so, does the loss function exhibit a benign optimization landscape? In this work, we provide partially affirmative answers to both questions for 1-hidden-layer networks with fewer than (sample size) neurons when the activation is smooth. First, we prove that as long as the width (where is the input dimension), its expressivity is strong, i.e., there exists at least one global minimizer with zero training loss. Second, we identify a nice local region with no local-min or saddle points. Nevertheless, it is not clear whether gradient descent can stay in this nice region. Third, we consider a constrained optimization formulation where the feasible region is the nice local region, and prove that every KKT point is a nearly global minimizer. It is expected that projected gradient methods converge to KKT points under mild technical conditions, but we leave the rigorous convergence analysis to future work. Thorough numerical results show that projected gradient methods on this constrained formulation significantly outperform SGD for training narrow neural nets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Neural Collapse with Normalized Features: A Geometric Analysis over the Riemannian ManifoldCan Yaras, Peng Wang, Zhihui Zhu, Laura Balzano 等NeurIPS 2022 · 被引用 60 次
- Why Robust Generalization in Deep Learning is Difficult: Perspective of Expressive PowerBinghui Li, Jikai Jin, Han Zhong, John E. Hopcroft 等NeurIPS 2022 · 被引用 37 次
- Memorization Capacity of Multi-Head Attention in TransformersSadegh Mahdavi, Renjie Liao, Christos ThrampoulidisICLR 2024 · 被引用 34 次
- Bounding the Width of Neural Networks via Coupled Initialization A Worst Case AnalysisAlexander Munteanu, Simon Omlor, Zhao Song, David P. WoodruffICML 2022 · 被引用 17 次
- pTNAS: Progressive Neural Architecture Search for Tabular DataNaili Xing, Shaofeng Cai, Lingze Zeng, Jiaqi Zhu 等ICML 2026 · 被引用 4 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 被引用 193 次
- Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization GuaranteeWei Hu, Zhiyuan Li, Dingli YuICLR 2020 · 被引用 140 次
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 被引用 128 次
- Perturbing Across the Feature Hierarchy to Improve Standard and Strict Blackbox Attack TransferabilityNathan Inkawhich, Kevin J. Liang, Binghui Wang, Matthew Inkawhich 等NeurIPS 2020 · 被引用 105 次
相关 Paper
- Beyond NTK with Vanilla Gradient Descent: A Mean-Field Analysis of Neural Networks with Polynomial Width, Samples, and TimeArvind V. Mahankali, Haochen Zhang, Kefan Dong, Margalit Glasgow 等NeurIPS 2023 · 被引用 20 次
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 被引用 41 次
- Spurious Valleys and Clustering Behavior of Neural NetworksSamuele PollaciICML 2023 · 被引用 1 次
- Optimization and Generalization of Shallow Neural Networks with Quadratic Activation FunctionsStefano Sarao Mannelli, Eric Vanden-Eijnden, Lenka ZdeborováNeurIPS 2020 · 被引用 65 次
- Bounds on Over-Parameterization for Guaranteed Existence of Descent Paths in Shallow ReLU NetworksArsalan Sharif-Nassab, Saber Salehkaleybar, S. Jamaloddin GolestaniICLR 2020 · 被引用 12 次
