Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods
Laurence Aitchison
摘要
We formulate the problem of neural network optimization as Bayesian filtering, where the observations are the backpropagated gradients. While neural network optimization has previously been studied using natural gradient methods which are closely related to Bayesian inference, they were unable to recover standard optimizers such as Adam and RMSprop with a root-mean-square gradient normalizer, instead getting a mean-square normalizer. To recover the root-mean-square normalizer, we find it necessary to account for the temporal dynamics of all the other parameters as they are geing optimized. The resulting optimizer, AdaBayes, adaptively transitions between SGD-like and Adam-like behaviour, automatically recovers AdamW, a state of the art variant of Adam with decoupled weight decay, and has generalisation performance competitive with SGD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 被引用 195 次
- Gradient Descent on Neurons and its Link to Approximate Second-order OptimizationFrederik BenzingICML 2022 · 被引用 31 次
- Hebbian Deep Learning Without FeedbackAdrien Journé, Hector Garcia Rodriguez, Qinghai Guo, Timoleon MoraitisICLR 2023 · 被引用 17 次
- Implicit Maximum a Posteriori Filtering via Adaptive OptimizationGianluca M. Bencomo, Jake Snell, Thomas L. GriffithsICLR 2024 · 被引用 4 次
- Robustness to corruption in pre-trained Bayesian neural networksXi Wang, Laurence AitchisonICLR 2023
相关 Paper
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 被引用 39 次
- Dynamic Momentum Recalibration in Online Gradient LearningZhipeng Yao, Rui Yu, Guisong Chang, Ying Li 等CVPR 2026 · 被引用 1 次
- Understanding Decoupled and Early Weight DecayJohan Bjorck, Kilian Q. Weinberger, Carla P. GomesAAAI 2021 · 被引用 37 次
- Learning in temporally structured environmentsMatt Jones, Tyler R. Scott, Mengye Ren, Gamaleldin Fathy Elsayed 等ICLR 2023 · 被引用 1 次
- FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large ModelsJunkang Liu, Fanhua Shang, Hongying Liu, Yuxuan Tian 等AAAI 2026 · 被引用 12 次
