How Tempering Fixes Data Augmentation in Bayesian Neural Networks
Gregor Bachmann, Lorenzo Noci, Thomas Hofmann
摘要
While Bayesian neural networks (BNNs) provide a sound and principled alternative to standard neural networks, an artificial sharpening of the posterior usually needs to be applied to reach comparable performance. This is in stark contrast to theory, dictating that given an adequate prior and a well-specified model, the untempered Bayesian posterior should achieve optimal performance. Despite the community's extensive efforts, the observed gains in performance still remain disputed with several plausible causes pointing at its origin. While data augmentation has been empirically recognized as one of the main drivers of this effect, a theoretical account of its role, on the other hand, is largely missing. In this work we identify two interlaced factors concurrently influencing the strength of the cold posterior effect, namely the correlated nature of augmentations and the degree of invariance of the employed model to such transformations. By theoretically analyzing simplified settings, we prove that tempering implicitly reduces the misspecification arising from modeling augmentations as i.i.d. data. The temperature mimics the role of the effective sample size, reflecting the gain in information provided by the augmentations. We corroborate our theoretical findings with extensive empirical evaluations, scaling to realistic BNNs. By relying on the framework of group convolutions, we experiment with models of varying inherent degree of invariance, confirming its hypothesized relationship with the optimal temperature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Optimizing Data Augmentation through Bayesian Model SelectionMadi Matymov, Ba-Hien Tran, Michael Kampffmeyer, Markus Heinonen 等ICLR 2026 · 被引用 1 次
- Gaussian Mean Field Variational Inference can Overestimate Predictive VarianceJames Odgers, Ben Riegler, Siddharth Swaroop, Vincent FortuinICML 2026
它引用的顶会 Paper13
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 被引用 1,747 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 被引用 458 次
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski 等ICML 2020 · 被引用 409 次
相关 Paper
- Disentangling the Roles of Curation, Data-Augmentation and the Prior in the Cold Posterior EffectLorenzo Noci, Kevin Roth, Gregor Bachmann, Sebastian Nowozin 等NeurIPS 2021 · 被引用 34 次
- On Uncertainty, Tempering, and Data Augmentation in Bayesian ClassificationSanyam Kapoor, Wesley J. Maddox, Pavel Izmailov, Andrew Gordon WilsonNeurIPS 2022 · 被引用 64 次
- A statistical theory of cold posteriors in deep neural networksLaurence AitchisonICLR 2021 · 被引用 7 次
- Bayesian Neural Network Priors RevisitedVincent Fortuin, Adrià Garriga-Alonso, Sebastian W. Ober, Florian Wenzel 等ICLR 2022 · 被引用 162 次
- Flat Seeking Bayesian Neural NetworksVan-Anh Nguyen, Tung-Long Vuong, Hoang Phan, Thanh-Toan Do 等NeurIPS 2023 · 被引用 14 次
