Scale-invariant Bayesian Neural Networks with Connectivity Tangent Kernel
Sungyub Kim, Sihwan Park, Kyung-Su Kim, Eunho Yang
摘要
Explaining generalizations and preventing over-confident predictions are central goals of studies on the loss landscape of neural networks. Flatness, defined as loss invariability on perturbations of a pre-trained solution, is widely accepted as a predictor of generalization in this context. However, the problem that flatness and generalization bounds can be changed arbitrarily according to the scale of a parameter was pointed out, and previous studies partially solved the problem with restrictions: Counter-intuitively, their generalization bounds were still variant for the function-preserving parameter scaling transformation or limited only to an impractical network structure. As a more fundamental solution, we propose new prior and posterior distributions invariant to scaling transformations by decomposing the scale and connectivity of parameters, thereby allowing the resulting generalization bound to describe the generalizability of a broad class of networks with the more practical class of transformations such as weight decay with batch normalization. We also show that the above issue adversely affects the uncertainty calibration of Laplace approximation and propose a solution using our invariant posterior. We empirically demonstrate our posterior provides effective flatness and calibration measures with low complexity in such a practical parameter transformation case, supporting its practical effectiveness in line with our rationale.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- The Geometry of Neural Nets' Parameter Spaces Under ReparametrizationAgustinus Kristiadi, Felix Dangel, Philipp HennigNeurIPS 2023 · 被引用 21 次
- Reparameterization invariance in approximate Bayesian inferenceHrittik Roy, Marco Miani, Carl Henrik Ek, Philipp Hennig 等NeurIPS 2024 · 被引用 20 次
- Stochastic Marginal Likelihood Gradients using Neural Tangent KernelsAlexander Immer, Tycho F. A. van der Ouderaa, Mark van der Wilk, Gunnar Rätsch 等ICML 2023 · 被引用 17 次
- Determinant Estimation under Memory Constraints and Neural Scaling LawsSiavash Ameli, Chris van der Heide, Liam Hodgkinson, Fred Roosta 等ICML 2025
- Spectral Estimation with Free DecompressionSiavash Ameli, Chris van der Heide, Liam Hodgkinson, Michael W. MahoneyNeurIPS 2025
它引用的顶会 Paper16
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 被引用 569 次
- Uncertainty Estimation Using a Single Deep Deterministic Neural NetworkJoost van Amersfoort, Lewis Smith, Yee Whye Teh, Yarin GalICML 2020 · 被引用 529 次
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
相关 Paper
- Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks Using PAC-Bayesian AnalysisYusuke Tsuzuku, Issei Sato, Masashi SugiyamaICML 2020 · 被引用 91 次
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 被引用 111 次
- Stable Minima of ReLU Neural Networks Suffer from the Curse of Dimensionality: The Neural Shattering PhenomenonTongtong Liang, Dan Qiao, Yu-Xiang Wang, Rahul ParhiNeurIPS 2025 · 被引用 8 次
- A Reparametrization-Invariant Sharpness Measure Based on Information GeometryCheongjae Jang, Sungyoon Lee, Frank C. Park, Yung-Kyun NohNeurIPS 2022 · 被引用 19 次
- Information-Theoretic Local Minima Characterization and RegularizationZhiwei Jia, Hao SuICML 2020 · 被引用 22 次
