The Sample Complexity of One-Hidden-Layer Neural Networks
Gal Vardi, Ohad Shamir, Nati Srebro
摘要
We study norm-based uniform convergence bounds for neural networks, aiming at a tight understanding of how these are affected by the architecture and type of norm constraint, for the simple class of scalar-valued one-hidden-layer networks, and inputs bounded in Euclidean norm. We begin by proving that in general, controlling the spectral norm of the hidden layer weight matrix is insufficient to get uniform convergence guarantees (independent of the network width), while a stronger Frobenius norm control is sufficient, extending and improving on previous work. Motivated by the proof constructions, we identify and analyze two important settings where (perhaps surprisingly) a mere spectral norm control turns out to be sufficient: First, when the network's activation functions are sufficiently smooth (with the result extending to deeper networks); and second, for certain types of convolutional networks. In the latter setting, we study how the sample complexity is additionally affected by parameters such as the amount of overlap between patches and the overall number of patches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- A PAC-Bayesian Generalization Bound for Equivariant NetworksArash Behboodi, Gabriele Cesa, Taco S. CohenNeurIPS 2022 · 被引用 26 次
- Generalization Bound and New Algorithm for Clean-Label Backdoor AttackLijia Yu, Shuang Liu, Yibo Miao, Xiao-Shan Gao 等ICML 2024 · 被引用 13 次
- Role of Locality and Weight Sharing in Image-Based Tasks: A Sample Complexity Separation between CNNs, LCNs, and FCNsAakash Lahoti, Stefani Karp, Ezra Winston, Aarti Singh 等ICLR 2024 · 被引用 5 次
- Initialization-Dependent Sample Complexity of Linear Predictors and Neural NetworksRoey Magen, Ohad ShamirNeurIPS 2023 · 被引用 2 次
- Generalizability of Neural Networks Minimizing Empirical Risk Based on Expressive PowerLijia Yu, Yibo Miao, Yifan Zhu, Xiao-Shan Gao 等ICLR 2025
它引用的顶会 Paper6
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Generalization bounds for deep convolutional neural networksPhilip M. Long, Hanie SedghiICLR 2020 · 被引用 102 次
- In Defense of Uniform Convergence: Generalization via Derandomization with an Application to Interpolating PredictorsJeffrey Negrea, Gintare Karolina Dziugaite, Daniel M. RoyICML 2020 · 被引用 66 次
- Norm-Based Generalisation Bounds for Deep Multi-Class Convolutional Neural NetworksAntoine Ledent, Waleed Mustafa, Yunwen Lei, Marius KloftAAAI 2021 · 被引用 24 次
相关 Paper
- Koopman-based generalization bound: New aspect for full-rank weightsYuka Hashimoto, Sho Sonoda, Isao Ishikawa, Atsushi Nitanda 等ICLR 2024 · 被引用 6 次
- Norm-based Generalization Bounds for Sparse Neural NetworksTomer Galanti, Mengjia Xu, Liane Galanti, Tomaso A. PoggioNeurIPS 2023 · 被引用 19 次
- Learning ReLU networks to high uniform accuracy is intractableJulius Berner, Philipp Grohs, Felix VoigtländerICLR 2023 · 被引用 2 次
- A closer look at the approximation capabilities of neural networksKai Fong Ernest ChongICLR 2020 · 被引用 18 次
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 被引用 172 次
