Flat Seeking Bayesian Neural Networks
Van-Anh Nguyen, Tung-Long Vuong, Hoang Phan, Thanh-Toan Do, Dinh Q. Phung, Trung Le
Abstract
Bayesian Neural Networks (BNNs) provide a probabilistic interpretation for deep learning models by imposing a prior distribution over model parameters and inferring a posterior distribution based on observed data. The model sampled from the posterior distribution can be used for providing ensemble predictions and quantifying prediction uncertainty. It is well-known that deep learning models with lower sharpness have better generalization ability. However, existing posterior inferences are not aware of sharpness/flatness in terms of formulation, possibly leading to high sharpness for the models sampled from them. In this paper, we develop theories, the Bayesian setting, and the variational inference approach for the sharpness-aware posterior. Specifically, the models sampled from our sharpness-aware posterior, and the optimal approximate posterior estimating this sharpness-aware posterior, have better flatness, hence possibly possessing higher generalization ability. We conduct experiments by leveraging the sharpness-aware posterior with state-of-the-art Bayesian Neural Networks, showing that the flat-seeking counterparts outperform their baselines in all metrics of interest.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0c569b9-d7e6-4f78-8f18-4b58cb740c51Cited by top-tier papers4
- Unveiling m-Sharpness Through the Structure of Stochastic Gradient NoiseHaocheng Luo, Mehrtash Harandi, Dinh Phung, Trung LeNeurIPS 2025 · 2 citations
- Promoting Ensemble Diversity with Interactive Bayesian Distributional Robustness for Fine-tuning Foundation ModelsNgoc-Quan Pham, Tuan Truong, Quyen Tran, Tan Minh Nguyen et al.ICML 2025
- Improving Generalization with Flat Hilbert Bayesian InferenceTuan Truong, Quyen Tran, Ngoc-Quan Pham, Nhat Ho et al.ICML 2025
- Flatness-Aware Stochastic Gradient Langevin DynamicsStefano Bruno, Youngsik Hwang, JaeHyeon An, Sotirios Sabanis et al.ICML 2026
Builds on22
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- SWAD: Domain Generalization by Seeking Flat MinimaJunbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho et al.NeurIPS 2021 · 630 citations
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 388 citations
Related papers
- Efficient and Scalable Bayesian Neural Nets with Rank-1 FactorsMichael Dusenberry, Ghassen Jerfel, Yeming Wen, Yi-An Ma et al.ICML 2020 · 239 citations
- Masked Bayesian Neural Networks : Theoretical Guarantee and its Posterior InferenceInsung Kong, Dongyoon Yang, Jongjin Lee, Ilsang Ohn et al.ICML 2023 · 8 citations
- Entropy-MCMC: Sampling from Flat Basins with EaseBolian Li, Ruqi ZhangICLR 2024 · 7 citations
- Microcanonical Langevin Ensembles: Advancing the Sampling of Bayesian Neural NetworksEmanuel Sommer, Jakob Robnik, Giorgi Nozadze, Uros Seljak et al.ICLR 2025
- On the Expressiveness of Approximate Inference in Bayesian Neural NetworksAndrew Y. K. Foong, David R. Burt, Yingzhen Li, Richard E. TurnerNeurIPS 2020 · 142 citations
