Mitigating Shortcut Learning with InterpoLated Learning
Michalis Korakakis, Andreas Vlachos, Adrian Weller
Abstract
Empirical risk minimization (ERM) incentivizes models to exploit shortcuts, i.e., spurious correlations between input attributes and labels that are prevalent in the majority of the training data but unrelated to the task at hand. This reliance hinders generalization on minority examples, where such correlations do not hold. Existing shortcut mitigation approaches are model-specific, difficult to tune, computationally expensive, and fail to improve learned representations. To address these issues, we propose InterpoLated Learning (Inter-poLL) which interpolates the representations of majority examples to include features from intra-class minority examples with shortcutmitigating patterns. This weakens shortcut influence, enabling models to acquire features predictive across both minority and majority examples. Experimental results on multiple natural language understanding tasks demonstrate that InterpoLL improves minority generalization over both ERM and state-of-the-art shortcut mitigation methods, without compromising accuracy on majority examples. Notably, these gains persist across encoder, encoder-decoder, and decoder-only architectures, demonstrating the method's broad applicability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb01509e-e1c3-4dee-a439-d744158a5379Builds on30
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 436 citations
- Improving Out-of-Distribution Robustness via Selective AugmentationHuaxiu Yao, Yu Wang, Sai Li, Linjun Zhang et al.ICML 2022 · 275 citations
Related papers
- COMI: COrrect and MItigate Shortcut Learning Behavior in Deep Neural NetworksLili Zhao, Qi Liu, Linan Yue, Wei Chen et al.SIGIR 2024 · 9 citations
- Which Shortcut Solution Do Question Answering Models Prefer to Learn?Kazutoshi Shinoda, Saku Sugawara, Akiko AizawaAAAI 2023 · 10 citations
- Gradient Extrapolation for Debiased Representation LearningIhab Asaad, Maha Shadaydeh, Joachim DenzlerICCV 2025 · 4 citations
- Improving the robustness of NLI models with minimax trainingMichalis Korakakis, Andreas VlachosACL 2023 · 4 citations
- Learning Concept Credible Models for Mitigating ShortcutsJiaxuan Wang, Sarah Jabbour, Maggie Makar, Michael W. Sjoding et al.NeurIPS 2022 · 8 citations
