When Expressivity Meets Trainability: Fewer than Neurons Can Work
Jiawei Zhang, Yushun Zhang, Mingyi Hong, Ruoyu Sun, Zhi-Quan Luo
Abstract
Modern neural networks are often quite wide, causing large memory and computation costs. It is thus of great interest to train a narrower network. However, training narrow neural nets remains a challenging task. We ask two theoretical questions: Can narrow networks have as strong expressivity as wide ones? If so, does the loss function exhibit a benign optimization landscape? In this work, we provide partially affirmative answers to both questions for 1-hidden-layer networks with fewer than (sample size) neurons when the activation is smooth. First, we prove that as long as the width (where is the input dimension), its expressivity is strong, i.e., there exists at least one global minimizer with zero training loss. Second, we identify a nice local region with no local-min or saddle points. Nevertheless, it is not clear whether gradient descent can stay in this nice region. Third, we consider a constrained optimization formulation where the feasible region is the nice local region, and prove that every KKT point is a nearly global minimizer. It is expected that projected gradient methods converge to KKT points under mild technical conditions, but we leave the rigorous convergence analysis to future work. Thorough numerical results show that projected gradient methods on this constrained formulation significantly outperform SGD for training narrow neural nets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Neural Collapse with Normalized Features: A Geometric Analysis over the Riemannian ManifoldCan Yaras, Peng Wang, Zhihui Zhu, Laura Balzano et al.NeurIPS 2022 · 60 citations
- Why Robust Generalization in Deep Learning is Difficult: Perspective of Expressive PowerBinghui Li, Jikai Jin, Han Zhong, John E. Hopcroft et al.NeurIPS 2022 · 37 citations
- Memorization Capacity of Multi-Head Attention in TransformersSadegh Mahdavi, Renjie Liao, Christos ThrampoulidisICLR 2024 · 34 citations
- Bounding the Width of Neural Networks via Coupled Initialization A Worst Case AnalysisAlexander Munteanu, Simon Omlor, Zhao Song, David P. WoodruffICML 2022 · 17 citations
- pTNAS: Progressive Neural Architecture Search for Tabular DataNaili Xing, Shaofeng Cai, Lingze Zeng, Jiaqi Zhu et al.ICML 2026 · 4 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 193 citations
- Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization GuaranteeWei Hu, Zhiyuan Li, Dingli YuICLR 2020 · 140 citations
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 128 citations
- Perturbing Across the Feature Hierarchy to Improve Standard and Strict Blackbox Attack TransferabilityNathan Inkawhich, Kevin J. Liang, Binghui Wang, Matthew Inkawhich et al.NeurIPS 2020 · 105 citations
Related papers
- Beyond NTK with Vanilla Gradient Descent: A Mean-Field Analysis of Neural Networks with Polynomial Width, Samples, and TimeArvind V. Mahankali, Haochen Zhang, Kefan Dong, Margalit Glasgow et al.NeurIPS 2023 · 20 citations
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 41 citations
- Spurious Valleys and Clustering Behavior of Neural NetworksSamuele PollaciICML 2023 · 1 citation
- Optimization and Generalization of Shallow Neural Networks with Quadratic Activation FunctionsStefano Sarao Mannelli, Eric Vanden-Eijnden, Lenka ZdeborováNeurIPS 2020 · 65 citations
- Bounds on Over-Parameterization for Guaranteed Existence of Descent Paths in Shallow ReLU NetworksArsalan Sharif-Nassab, Saber Salehkaleybar, S. Jamaloddin GolestaniICLR 2020 · 12 citations
