Bayesian Invariant Risk Minimization
Yong Lin, Hanze Dong, Hao Wang, Tong Zhang
Abstract
Generalization under distributional shift is an open challenge for machine learning. Invariant Risk Minimization (IRM) is a promising framework to tackle this issue by extracting invariant features. However, despite the potential and popularity of IRM, recent works have reported negative results of it on deep models. We argue that the failure can be primarily attributed to deep models' tendency to overfit the data. Specifically, our theoretical analysis shows that IRM degenerates to empirical risk minimization (ERM) when overfitting occurs. Our empirical evidence also provides supports: IRM methods that work well in typical settings significantly deteriorate even if we slightly enlarge the model size or lessen the training data. To alleviate this issue, we propose Bayesian Invariant Risk Min-imization (BIRM) by introducing Bayesian inference into the IRM. The key motivation is to estimate the penalty of IRM based on the posterior distribution of classifiers (as opposed to a single classifier), which is much less prone to overfitting. Extensive experimental results on four datasets demonstrate that BIRM consistently outperforms the existing IRM baselines significantly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03836bbe-bfee-47c7-b6a4-aa31705ca862Cited by top-tier papers29
- Learning Causally Invariant Representations for Out-of-Distribution Generalization on GraphsYongqiang Chen, Yonggang Zhang, Yatao Bian, Han Yang et al.NeurIPS 2022 · 246 citations
- ZIN: When and How to Learn Invariance Without Environment Partition?Yong Lin, Shengyu Zhu, Lu Tan, Peng CuiNeurIPS 2022 · 91 citations
- Sparse Invariant Risk MinimizationXiao Zhou, Yong Lin, Weizhong Zhang, Tong ZhangICML 2022 · 85 citations
- Model Agnostic Sample Reweighting for Out-of-Distribution LearningXiao Zhou, Yong Lin, Renjie Pi, Weizhong Zhang et al.ICML 2022 · 73 citations
- Bayesian Prompt Learning for Image-Language Model GeneralizationMohammad Mahdi Derakhshani, Enrique Sanchez, Adrian Bulat, Victor Guilherme Turrisi da Costa et al.ICCV 2023 · 66 citations
Builds on19
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain et al.NeurIPS 2020 · 503 citations
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 454 citations
- The Risks of Invariant Risk MinimizationElan Rosenfeld, Pradeep Kumar Ravikumar, Andrej RisteskiICLR 2021 · 356 citations
Related papers
- Learning Optimal Features via Partial InvarianceMoulik Choraria, Ibtihal Ferwana, Ankur Mani, Lav R. VarshneyAAAI 2023 · 3 citations
- Empirical or Invariant Risk Minimization? A Sample Complexity PerspectiveKartik Ahuja, Jun Wang, Amit Dhurandhar, Karthikeyan Shanmugam et al.ICLR 2021 · 15 citations
- Distribution Shift Is Key to Learning Invariant PredictionHong Zheng, Fei TengAAAI 2026
- IRM - when it works and when it doesn't: A test case of natural language inferenceYana Dranker, He He, Yonatan BelinkovNeurIPS 2021 · 22 citations
- On the Connection between Invariant Learning and Adversarial Training for Out-of-Distribution GeneralizationShiji Xin, Yifei Wang, Jingtong Su, Yisen WangAAAI 2023 · 14 citations
