Non-Vacuous Generalisation Bounds for Shallow Neural Networks
Felix Biggs, Benjamin Guedj
Abstract
We focus on a specific class of shallow neural networks with a single hidden layer, namely those with L 2 -normalised data and either a sigmoidshaped Gaussian error function ("erf") activation or a Gaussian Error Linear Unit (GELU) activation. For these networks, we derive new generalisation bounds through the PAC-Bayesian theory; unlike most existing such bounds they apply to neural networks with deterministic rather than randomised parameters. Our bounds are empirically non-vacuous when the network is trained with vanilla stochastic gradient descent on MNIST, Fashion-MNIST, and binary classification versions of the above.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f322a39-2f8a-46cd-8884-fd6dba6c1411Cited by top-tier papers5
- MMD-Fuse: Learning and Combining Kernels for Two-Sample Testing Without Data SplittingFelix Biggs, Antonin Schrab, Arthur GrettonNeurIPS 2023 · 49 citations
- Learning via Wasserstein-Based High Probability Generalisation BoundsPaul Viallard, Maxime Haddouche, Umut Simsekli, Benjamin GuedjNeurIPS 2023 · 16 citations
- On Margins and Generalisation for Voting ClassifiersFelix Biggs, Valentina Zantedeschi, Benjamin GuedjNeurIPS 2022 · 10 citations
- Generalization Bounds for Rank-sparse Neural NetworksAntoine Ledent, Rodrigo Alves, Yunwen LeiNeurIPS 2025 · 4 citations
- Controlling Multiple Errors Simultaneously with a PAC-Bayes BoundReuben Adams, John Shawe-Taylor, Benjamin GuedjNeurIPS 2024
Builds on4
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- In search of robust measures of generalizationGintare Karolina Dziugaite, Alexandre Drouin, Brady Neal, Nitarshan Rajkumar et al.NeurIPS 2020 · 112 citations
- Second Order PAC-Bayesian Bounds for the Weighted Majority VoteAndrés R. Masegosa, Stephan Sloth Lorenzen, Christian Igel, Yevgeny SeldinNeurIPS 2020 · 48 citations
- Learning Stochastic Majority Votes by Minimizing a PAC-Bayes Generalization BoundValentina Zantedeschi, Paul Viallard, Emilie Morvant, Rémi Emonet et al.NeurIPS 2021 · 21 citations
Related papers
- Should Under-parameterized Student Networks Copy or Average Teacher Weights?Berfin Simsek, Amire Bendjeddou, Wulfram Gerstner, Johanni BreaNeurIPS 2023 · 14 citations
- On the Universality of the Double Descent Peak in Ridgeless RegressionDavid HolzmüllerICLR 2021 · 16 citations
- Agnostic Learning of a Single Neuron with Gradient DescentSpencer Frei, Yuan Cao, Quanquan GuNeurIPS 2020 · 68 citations
- Sharp Generalization for Nonparametric Regression by Over-Parameterized Neural Networks: A Distribution-Free Analysis in Spherical CovariateYingzhen YangICML 2025
- Avoiding Kernel Fixed Points: Computing with ELU and GELU Infinite NetworksRussell Tsuchida, Tim Pearce, Christopher van der Heide, Fred Roosta et al.AAAI 2021 · 10 citations
