Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets
Yuxiang Wu, Matt Gardner, Pontus Stenetorp, Pradeep Dasigi
Abstract
Natural language processing models often exploit spurious correlations between task-independent features and labels in datasets to perform well only within the distributions they are trained on, while not generalising to different task distributions. We propose to tackle this problem by generating a debiased version of a dataset, which can then be used to train a debiased, off-the-shelf model, by simply replacing its training data. Our approach consists of 1) a method for training data generators to generate high-quality, label-consistent data samples; and 2) a filtering mechanism for removing data points that contribute to spurious correlations, measured in terms of z-statistics. We generate debiased versions of the SNLI and MNLI datasets, and we evaluate on a large suite of debiased, out-of-distribution, and adversarial test sets. Results show that models trained on our debiased datasets generalise better than those trained on the original datasets in all settings. On the majority of the datasets, our method outperforms or performs comparably to previous state-of-the-art debiasing strategies, and when combined with an orthogonal technique, product-of-experts, it improves further and outperforms previous best results of SNLI-hard and MNLI-hard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b581dfb1-0dc1-4e06-bce2-383e03db50e8Cited by top-tier papers20
- PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for FreeHao Li, Xiaogeng Liu, Ning Zhang, Chaowei XiaoACL 2025 · 33 citations
- DISCO: Distilling Counterfactuals with Large Language ModelsZeming Chen, Qiyue Gao, Antoine Bosselut, Ashish Sabharwal et al.ACL 2023 · 27 citations
- BITE: Textual Backdoor Attacks with Iterative Trigger InjectionJun Yan, Vansh Gupta, Xiang RenACL 2023 · 24 citations
- Feature-Level Debiased Natural Language UnderstandingYougang Lyu, Piji Li, Yechang Yang, Maarten de Rijke et al.AAAI 2023 · 12 citations
- Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureLinyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu et al.EMNLP 2023 · 12 citations
Builds on13
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers et al.ICML 2020 · 242 citations
- End-to-End Bias Mitigation by Modelling Biases in CorporaRabeeh Karimi Mahabadi, Yonatan Belinkov, James HendersonACL 2020 · 136 citations
- Learning from others' mistakes: Avoiding dataset biases without modeling themVictor Sanh, Thomas Wolf, Yonatan Belinkov, Alexander M. RushICLR 2021 · 123 citations
Related papers
- BiasAdv: Bias-Adversarial Augmentation for Model DebiasingJongin Lim, Youngdong Kim, Byungjai Kim, Chanho Ahn et al.CVPR 2023
- Towards Robustifying NLI Models Against Lexical Dataset BiasesXiang Zhou, Mohit BansalACL 2020 · 36 citations
- DeNetDM: Debiasing by Network Depth ModulationSilpa Vadakkeeveetil Sreelatha, Adarsh Kappiyath, Abhra Chaudhuri, Anjan DuttaNeurIPS 2024 · 8 citations
- Improving the robustness of NLI models with minimax trainingMichalis Korakakis, Andreas VlachosACL 2023 · 4 citations
- Echoes: Unsupervised Debiasing via Pseudo-bias Labeling in an Echo ChamberRui Hu, Yahan Tu, Jitao SangACM MM 2023 · 1 citation
