The Expected Loss of Preconditioned Langevin Dynamics Reveals the Hessian Rank
Amitay Bar, Rotem Mulayoff, Tomer Michaeli, Ronen Talmon
摘要
Langevin dynamics (LD) is widely used for sampling from distributions and for optimization. In this work, we derive a closed-form expression for the expected loss of preconditioned LD near stationary points of the objective function. We use the fact that at the vicinity of such points, LD reduces to an Ornstein–Uhlenbeck process, which is amenable to convenient mathematical treatment. Our analysis reveals that when the preconditioning matrix satisfies a particular relation with respect to the noise covariance, LD's expected loss becomes proportional to the rank of the objective's Hessian. We illustrate the applicability of this result in the context of neural networks, where the Hessian rank has been shown to capture the complexity of the predictor function but is usually computationally hard to probe. Finally, we use our analysis to compare SGD-like and Adam-like preconditioners and identify the regimes under which each of them leads to a lower expected loss.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 被引用 155 次
- Continuous vs. Discrete Optimization of Deep Neural NetworksOmer Elkabetz, Nadav CohenNeurIPS 2021 · 被引用 51 次
- Unique Properties of Flat Minima in Deep NetworksRotem Mulayoff, Tomer MichaeliICML 2020 · 被引用 43 次
相关 Paper
- How Does Adaptive Optimization Impact Local Neural Network Geometry?Kaiqi Jiang, Dhruv Malik, Yuanzhi LiNeurIPS 2023 · 被引用 26 次
- Adaptive Preconditioners Trigger Loss Spikes in AdamZhiwei Bai, Zhangchen Zhou, Jiajie Zhao, Xiaolong Li 等ICML 2026 · 被引用 9 次
- ASGO: Adaptive Structured Gradient OptimizationKang An, Yuxing Liu, Rui Pan, Yi Ren 等NeurIPS 2025 · 被引用 58 次
- Exact risk curves of signSGD in High-Dimensions: quantifying preconditioning and noise-compression effectsKe Liang Xiao, Noah Marshall, Atish Agarwala, Elliot PaquetteICML 2025
- AGD: an Auto-switchable Optimizer using Stepwise Gradient Difference for Preconditioning MatrixYun Yue, Zhiling Ye, Jiadi Jiang, Yongchao Liu 等NeurIPS 2023 · 被引用 6 次
