IRM - when it works and when it doesn't: A test case of natural language inference
Yana Dranker, He He, Yonatan Belinkov
Abstract
Invariant Risk Minimization (IRM) is a recently proposed framework for outof-distribution (o.o.d) generalization. Most of the studies on IRM so far have focused on theoretical results, toy problems, and simple models. In this work, we investigate the applicability of IRM to bias mitigation-a special case of o.o.d generalization-in increasingly naturalistic settings and deep models. Using natural language inference (NLI) as a test case, we start with a setting where both the dataset and the bias are synthetic, continue with a natural dataset and synthetic bias, and end with a fully realistic setting with natural datasets and bias. Our results show that in naturalistic settings, learning complex features in place of the bias proves to be difficult, leading to a rather small improvement over empirical risk minimization. Moreover, we find that in addition to being sensitive to random seeds, the performance of IRM also depends on several critical factors, notably dataset size, bias prevalence, and bias strength, thus limiting IRM's advantage in practical scenarios. Our results highlight key challenges in applying IRM to real-world scenarios, calling for a more naturalistic characterization of the problem setup for o.o.d generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eef4eb5f-7da0-4927-9719-2f523e065cb9Cited by top-tier papers6
- Are All Spurious Features in Natural Language Alike? An Analysis through a Causal LensNitish Joshi, Xiang Pan, He HeEMNLP 2022 · 19 citations
- Causal-structure Driven Augmentations for Text OOD GeneralizationAmir Feder, Yoav Wald, Claudia Shi, Suchi Saria et al.NeurIPS 2023 · 10 citations
- Environment Diversification with Multi-head Neural Network for Invariant LearningBo-Wei Huang, Keng-Te Liao, Chang-Sheng Kao, Shou-De LinNeurIPS 2022 · 5 citations
- What Is Missing in IRM Training and Evaluation? Challenges and SolutionsYihua Zhang, Pranay Sharma, Parikshit Ram, Mingyi Hong et al.ICLR 2023
- An Invariant Learning Characterization of Controlled Text GenerationCarolina Zheng, Claudia Shi, Keyon Vafa, Amir Feder et al.ACL 2023
Builds on8
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- The Risks of Invariant Risk MinimizationElan Rosenfeld, Pradeep Kumar Ravikumar, Andrej RisteskiICLR 2021 · 356 citations
- Invariant Risk Minimization GamesKartik Ahuja, Karthikeyan Shanmugam, Kush R. Varshney, Amit DhurandharICML 2020 · 289 citations
Related papers
- Bayesian Invariant Risk MinimizationYong Lin, Hanze Dong, Hao Wang, Tong ZhangCVPR 2022 · 48 citations
- Sparse Invariant Risk MinimizationXiao Zhou, Yong Lin, Weizhong Zhang, Tong ZhangICML 2022 · 85 citations
- Empirical or Invariant Risk Minimization? A Sample Complexity PerspectiveKartik Ahuja, Jun Wang, Amit Dhurandhar, Karthikeyan Shanmugam et al.ICLR 2021 · 15 citations
- Invariant Language ModelingMaxime Peyrard, Sarvjeet Singh Ghotra, Martin Josifoski, Vidhan Agarwal et al.EMNLP 2022 · 8 citations
- The Missing Invariance Principle found - the Reciprocal Twin of Invariant Risk MinimizationDongsung Huh, Avinash BaidyaNeurIPS 2022 · 11 citations
