Variational Learning Finds Flatter Solutions at the Edge of Stability
Avrajit Ghosh, Bai Cong, Rio Yokota, Saiprasad Ravishankar, Rongrong Wang, Molei Tao, Mohammad Emtiyaz Khan, Thomas Möllenhoff
摘要
Variational Learning (VL) has recently gained popularity for training deep neural networks. Part of its empirical success can be explained by theories such as PAC-Bayes bounds, minimum description length and marginal likelihood, but little has been done to unravel the implicit regularization in play. Here, we analyze the implicit regularization of VL through the Edge of Stability (EoS) framework. EoS has previously been used to show that gradient descent can find flat solutions and we extend this result to show that VL can find even flatter solutions. This result is obtained by controlling the shape of the variational posterior as well as the number of posterior samples used during training. The derivation follows in a similar fashion as in the standard EoS literature for deep learning, by first deriving a result for a quadratic problem and then extending it to deep neural networks. We empirically validate these findings on a wide variety of large networks, such as ResNet and ViT, to find that the theoretical results closely match the empirical ones. Ours is the first work to analyze the EoS dynamics of VL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit 等ICLR 2020 · 被引用 198 次
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 被引用 165 次
- Stabilizing Transformer Training by Preventing Attention Entropy CollapseShuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge 等ICML 2023 · 被引用 153 次
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 被引用 139 次
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch 等ICML 2021 · 被引用 130 次
相关 Paper
- Surprising Instabilities in Training Deep Networks and a Theoretical AnalysisYuxin Sun, Dong Lao, Ganesh Sundaramoorthi, Anthony J. YezziNeurIPS 2022 · 被引用 15 次
- A Minimalist Example of Edge-of-Stability and Progressive SharpeningLiming Liu, Zixuan Zhang, Simon S. Du, Tuo ZhaoNeurIPS 2025 · 被引用 4 次
- SAM operates far from home: eigenvalue regularization as a dynamical phenomenonAtish Agarwala, Yann N. DauphinICML 2023 · 被引用 26 次
- Variational Deep Learning via Implicit RegularizationJonathan Wenger, Beau Coker, Juraj Marusic, John Patrick CunninghamICLR 2026 · 被引用 1 次
- Understanding Edge-of-Stability Training Dynamics with a Minimalist ExampleXingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou 等ICLR 2023 · 被引用 1 次
