Variational Learning is Effective for Large Deep Networks
Yuesong Shen, Nico Daheim, Bai Cong, Peter Nickl, Gian Maria Marconi, Clement Bazan, Rio Yokota, Iryna Gurevych, Daniel Cremers, Mohammad Emtiyaz Khan, Thomas Möllenhoff
摘要
We give extensive empirical evidence against the common belief that variational learning is ineffective for large neural networks. We show that an optimizer called Improved Variational Online Newton (IVON) consistently matches or outperforms Adam for training large networks such as GPT-2 and ResNets from scratch. IVON's computational costs are nearly identical to Adam but its predictive uncertainty is better. We show several new use cases of IVON where we improve finetuning and model merging in Large Language Models, accurately predict generalization error, and faithfully estimate sensitivity to data. We find overwhelming evidence that variational learning is effective. Code is available at https://github.com/team-approx-bayes/ivon .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Bayesian Online Natural Gradient (BONG)Matt Jones, Peter G. Chang, Kevin P. MurphyNeurIPS 2024 · 被引用 20 次
- CLEAR: Calibrated Learning for Epistemic and Aleatoric RiskIlia Azizi, Juraj Bodik, Jakob Heiss, Bin YuICLR 2026 · 被引用 9 次
- Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement LearningMohamed Elsayed, Homayoon Farrahi, Felix Dangel, A. Rupam MahmoodICML 2024 · 被引用 7 次
- Federated ADMM from Bayesian DualityThomas Möllenhoff, Siddharth Swaroop, Finale Doshi-Velez, Mohammad Emtiyaz KhanICLR 2026 · 被引用 4 次
- Asymmetric Duos: Sidekicks Improve UncertaintyTim G. Zhou, Evan Shelhamer, Geoff PleissNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper14
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 被引用 458 次
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski 等ICML 2020 · 被引用 409 次
相关 Paper
- Gradient descent with generalized Newton's methodZhiqi Bu, Shiyun XuICLR 2025
- MARS: Unleashing the Power of Variance Reduction for Training Large ModelsHuizhuo Yuan, Yifeng Liu, Shuang Wu, Xun Zhou 等ICML 2025
- In Search of Adam's Secret SauceAntonio Orvieto, Robert GowerNeurIPS 2025 · 被引用 43 次
- A Physics-Inspired Optimizer: Velocity Regularized AdamPranav Vaidhyanathan, Lucas Schorling, Natalia Ares, Michael A. OsborneICLR 2026 · 被引用 1 次
- Structured Stochastic Gradient MCMCAntonios Alexos, Alex J. Boyd, Stephan MandtICML 2022 · 被引用 14 次
