Variational Learning is Effective for Large Deep Networks
Yuesong Shen, Nico Daheim, Bai Cong, Peter Nickl, Gian Maria Marconi, Clement Bazan, Rio Yokota, Iryna Gurevych, Daniel Cremers, Mohammad Emtiyaz Khan, Thomas Möllenhoff
Abstract
We give extensive empirical evidence against the common belief that variational learning is ineffective for large neural networks. We show that an optimizer called Improved Variational Online Newton (IVON) consistently matches or outperforms Adam for training large networks such as GPT-2 and ResNets from scratch. IVON's computational costs are nearly identical to Adam but its predictive uncertainty is better. We show several new use cases of IVON where we improve finetuning and model merging in Large Language Models, accurately predict generalization error, and faithfully estimate sensitivity to data. We find overwhelming evidence that variational learning is effective. Code is available at https://github.com/team-approx-bayes/ivon .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc34cb4b-57bf-4d54-a66a-17cc12de3f4fCited by top-tier papers22
- Bayesian Online Natural Gradient (BONG)Matt Jones, Peter G. Chang, Kevin P. MurphyNeurIPS 2024 · 20 citations
- CLEAR: Calibrated Learning for Epistemic and Aleatoric RiskIlia Azizi, Juraj Bodik, Jakob Heiss, Bin YuICLR 2026 · 9 citations
- Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement LearningMohamed Elsayed, Homayoon Farrahi, Felix Dangel, A. Rupam MahmoodICML 2024 · 7 citations
- Federated ADMM from Bayesian DualityThomas Möllenhoff, Siddharth Swaroop, Finale Doshi-Velez, Mohammad Emtiyaz KhanICLR 2026 · 4 citations
- Asymmetric Duos: Sidekicks Improve UncertaintyTim G. Zhou, Evan Shelhamer, Geoff PleissNeurIPS 2025 · 3 citations
Builds on14
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen et al.NeurIPS 2021 · 508 citations
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
Related papers
- Gradient descent with generalized Newton's methodZhiqi Bu, Shiyun XuICLR 2025
- MARS: Unleashing the Power of Variance Reduction for Training Large ModelsHuizhuo Yuan, Yifeng Liu, Shuang Wu, Xun Zhou et al.ICML 2025
- In Search of Adam's Secret SauceAntonio Orvieto, Robert GowerNeurIPS 2025 · 43 citations
- A Physics-Inspired Optimizer: Velocity Regularized AdamPranav Vaidhyanathan, Lucas Schorling, Natalia Ares, Michael A. OsborneICLR 2026 · 1 citation
- Structured Stochastic Gradient MCMCAntonios Alexos, Alex J. Boyd, Stephan MandtICML 2022 · 14 citations
