Variational Learning Finds Flatter Solutions at the Edge of Stability
Avrajit Ghosh, Bai Cong, Rio Yokota, Saiprasad Ravishankar, Rongrong Wang, Molei Tao, Mohammad Emtiyaz Khan, Thomas Möllenhoff
Abstract
Variational Learning (VL) has recently gained popularity for training deep neural networks. Part of its empirical success can be explained by theories such as PAC-Bayes bounds, minimum description length and marginal likelihood, but little has been done to unravel the implicit regularization in play. Here, we analyze the implicit regularization of VL through the Edge of Stability (EoS) framework. EoS has previously been used to show that gradient descent can find flat solutions and we extend this result to show that VL can find even flatter solutions. This result is obtained by controlling the shape of the variational posterior as well as the number of posterior samples used during training. The derivation follows in a similar fashion as in the standard EoS literature for deep learning, by first deriving a result for a quadratic problem and then extending it to deep neural networks. We empirically validate these findings on a wide variety of large networks, such as ResNet and ViT, to find that the theoretical results closely match the empirical ones. Ours is the first work to analyze the EoS dynamics of VL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on24
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 165 citations
- Stabilizing Transformer Training by Preventing Attention Entropy CollapseShuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge et al.ICML 2023 · 153 citations
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 139 citations
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch et al.ICML 2021 · 130 citations
Related papers
- Surprising Instabilities in Training Deep Networks and a Theoretical AnalysisYuxin Sun, Dong Lao, Ganesh Sundaramoorthi, Anthony J. YezziNeurIPS 2022 · 15 citations
- A Minimalist Example of Edge-of-Stability and Progressive SharpeningLiming Liu, Zixuan Zhang, Simon S. Du, Tuo ZhaoNeurIPS 2025 · 4 citations
- SAM operates far from home: eigenvalue regularization as a dynamical phenomenonAtish Agarwala, Yann N. DauphinICML 2023 · 26 citations
- Variational Deep Learning via Implicit RegularizationJonathan Wenger, Beau Coker, Juraj Marusic, John Patrick CunninghamICLR 2026 · 1 citation
- Understanding Edge-of-Stability Training Dynamics with a Minimalist ExampleXingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou et al.ICLR 2023 · 1 citation
