Lune

ICML2026Top-tier venue

Flatness-Aware Stochastic Gradient Langevin Dynamics

Stefano Bruno, Youngsik Hwang, JaeHyeon An, Sotirios Sabanis, Dongyoung Lim

2026Year

Abstract

Flatness of the loss landscape has been widely studied as an important perspective for understanding the behavior and generalization of deep learning algorithms. Motivated by this view, we propose Flatness-Aware Stochastic Gradient Langevin Dynamics (fSGLD), a first-order optimization method that biases learning its dynamics toward flat basins while retaining the computational and memory efficiency of SGD and SGLD. We provide a non-asymptotic theoretical analysis showing that fSGLD targets a flatness-biased Gibbs distribution under a theoretically prescribed coupling between the noise scale σ and the inverse temperature β, together with explicit excess risk guarantees. We empirically evaluate fSGLD across standard optimizer benchmarks, Bayesian image classification, uncertainty quantification, and out-of-distribution detection, demonstrating consistently strong performance and reliable uncertainty estimates. Additional experiments confirm the effectiveness of the theoretically prescribed β-σ coupling compared to decoupled choices. The code for all the experiments is available at https: //github.com/youngsikhwang/ Flatness-aware-SGLD.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c0e63f79-9b50-43ae-b8d9-ef63e4253d08

Builds on38

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines