On Energy-Based Models with Overparametrized Shallow Neural Networks
Carles Domingo-Enrich, Alberto Bietti, Eric Vanden-Eijnden, Joan Bruna
Abstract
Energy-based models (EBMs) are a simple yet powerful framework for generative modeling. They are based on a trainable energy function which defines an associated Gibbs measure, and they can be trained and sampled from via well-established statistical tools, such as MCMC. Neural networks may be used as energy function approximators, providing both a rich class of expressive models as well as a flexible device to incorporate data structure. In this work we focus on shallow neural networks. Building from the incipient theory of overparametrized neural networks, we show that models trained in the so-called "active" regime provide a statistical advantage over their associated "lazy" or kernel regime, leading to improved adaptivity to hidden low-dimensional structure in the data distribution, as already observed in supervised learning. Our study covers both maximum likelihood and Stein Discrepancy estimators, and we validate our theoretical results with numerical experiments on synthetic data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Local Signal Adaptivity: Provable Feature Learning in Neural Networks Beyond KernelsStefani Karp, Ezra Winston, Yuanzhi Li, Aarti SinghNeurIPS 2021 · 38 citations
- Efficient Training of Energy-Based Models Using Jarzynski EqualityDavide Carbone, Mengjian Hua, Simon Coste, Eric Vanden-EijndenNeurIPS 2023 · 21 citations
- Score-based generative models break the curse of dimensionality in learning a family of sub-Gaussian distributionsFrank Cole, Yulong LuICLR 2024 · 9 citations
- Conditionally Strongly Log-Concave Generative ModelsFlorentin Guth, Etienne Lempereur, Joan Bruna, Stéphane MallatICML 2023 · 5 citations
- Generalizable Reasoning through Compositional Energy MinimizationAlexandru Oarga, Yilun DuNeurIPS 2025 · 3 citations
Builds on4
- When Do Neural Networks Outperform Kernel Methods?Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, Andrea MontanariNeurIPS 2020 · 217 citations
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 172 citations
- Learning the Stein Discrepancy for Training and Evaluating Energy-Based Models without SamplingWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICML 2020 · 93 citations
- Quantifying the Benefit of Using Differentiable Learning over Tangent KernelsEran Malach, Pritish Kamath, Emmanuel Abbe, Nathan SrebroICML 2021 · 44 citations
Related papers
- Training Deep Energy-Based Models with f-Divergence MinimizationLantao Yu, Yang Song, Jiaming Song, Stefano ErmonICML 2020 · 50 citations
- No MCMC for me: Amortized sampling for fast and stable training of energy-based modelsWill Sussman Grathwohl, Jacob Jin Kelly, Milad Hashemi, Mohammad Norouzi et al.ICLR 2021 · 75 citations
- Explaining the effects of non-convergent MCMC in the training of Energy-Based ModelsElisabeth Agoritsas, Giovanni Catania, Aurélien Decelle, Beatriz SeoaneICML 2023 · 17 citations
- Deep Equals Shallow for ReLU Networks in Kernel RegimesAlberto Bietti, Francis R. BachICLR 2021 · 9 citations
- Hamiltonian Dynamics with Non-Newtonian Momentum for Rapid SamplingGreg Ver Steeg, Aram GalstyanNeurIPS 2021 · 18 citations
