The Expected Loss of Preconditioned Langevin Dynamics Reveals the Hessian Rank
Amitay Bar, Rotem Mulayoff, Tomer Michaeli, Ronen Talmon
Abstract
Langevin dynamics (LD) is widely used for sampling from distributions and for optimization. In this work, we derive a closed-form expression for the expected loss of preconditioned LD near stationary points of the objective function. We use the fact that at the vicinity of such points, LD reduces to an Ornstein–Uhlenbeck process, which is amenable to convenient mathematical treatment. Our analysis reveals that when the preconditioning matrix satisfies a particular relation with respect to the noise covariance, LD's expected loss becomes proportional to the rank of the objective's Hessian. We illustrate the applicability of this result in the context of neural networks, where the Hessian rank has been shown to capture the complexity of the predictor function but is usually computationally hard to probe. Finally, we use our analysis to compare SGD-like and Adam-like preconditioners and identify the regimes under which each of them leads to a lower expected loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 155 citations
- Continuous vs. Discrete Optimization of Deep Neural NetworksOmer Elkabetz, Nadav CohenNeurIPS 2021 · 51 citations
- Unique Properties of Flat Minima in Deep NetworksRotem Mulayoff, Tomer MichaeliICML 2020 · 43 citations
Related papers
- How Does Adaptive Optimization Impact Local Neural Network Geometry?Kaiqi Jiang, Dhruv Malik, Yuanzhi LiNeurIPS 2023 · 26 citations
- Adaptive Preconditioners Trigger Loss Spikes in AdamZhiwei Bai, Zhangchen Zhou, Jiajie Zhao, Xiaolong Li et al.ICML 2026 · 9 citations
- ASGO: Adaptive Structured Gradient OptimizationKang An, Yuxing Liu, Rui Pan, Yi Ren et al.NeurIPS 2025 · 58 citations
- Exact risk curves of signSGD in High-Dimensions: quantifying preconditioning and noise-compression effectsKe Liang Xiao, Noah Marshall, Atish Agarwala, Elliot PaquetteICML 2025
- AGD: an Auto-switchable Optimizer using Stepwise Gradient Difference for Preconditioning MatrixYun Yue, Zhiling Ye, Jiadi Jiang, Yongchao Liu et al.NeurIPS 2023 · 6 citations
