Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods
Laurence Aitchison
Abstract
We formulate the problem of neural network optimization as Bayesian filtering, where the observations are the backpropagated gradients. While neural network optimization has previously been studied using natural gradient methods which are closely related to Bayesian inference, they were unable to recover standard optimizers such as Adam and RMSprop with a root-mean-square gradient normalizer, instead getting a mean-square normalizer. To recover the root-mean-square normalizer, we find it necessary to account for the temporal dynamics of all the other parameters as they are geing optimized. The resulting optimizer, AdaBayes, adaptively transitions between SGD-like and Adam-like behaviour, automatically recovers AdamW, a state of the art variant of Adam with decoupled weight decay, and has generalisation performance competitive with SGD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 195 citations
- Gradient Descent on Neurons and its Link to Approximate Second-order OptimizationFrederik BenzingICML 2022 · 31 citations
- Hebbian Deep Learning Without FeedbackAdrien Journé, Hector Garcia Rodriguez, Qinghai Guo, Timoleon MoraitisICLR 2023 · 17 citations
- Implicit Maximum a Posteriori Filtering via Adaptive OptimizationGianluca M. Bencomo, Jake Snell, Thomas L. GriffithsICLR 2024 · 4 citations
- Robustness to corruption in pre-trained Bayesian neural networksXi Wang, Laurence AitchisonICLR 2023
Related papers
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 39 citations
- Dynamic Momentum Recalibration in Online Gradient LearningZhipeng Yao, Rui Yu, Guisong Chang, Ying Li et al.CVPR 2026 · 1 citation
- Understanding Decoupled and Early Weight DecayJohan Bjorck, Kilian Q. Weinberger, Carla P. GomesAAAI 2021 · 37 citations
- Learning in temporally structured environmentsMatt Jones, Tyler R. Scott, Mengye Ren, Gamaleldin Fathy Elsayed et al.ICLR 2023 · 1 citation
- FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large ModelsJunkang Liu, Fanhua Shang, Hongying Liu, Yuxuan Tian et al.AAAI 2026 · 12 citations
