Towards Robustifying NLI Models Against Lexical Dataset Biases
Xiang Zhou, Mohit Bansal
Abstract
While deep learning models are making fast progress on the task of Natural Language Inference, recent studies have also shown that these models achieve high accuracy by exploiting several dataset biases, and without deep understanding of the language semantics. Using contradiction-word bias and word-overlapping bias as our two bias examples, this paper explores both data-level and model-level debiasing methods to robustify models against lexical dataset biases. First, we debias the dataset through data augmentation and enhancement, but show that the model bias cannot be fully removed via this method. Next, we also compare two ways of directly debiasing the model without knowing what the dataset biases are in advance. The first approach aims to remove the label bias at the embedding level. The second approach employs a bag-of-words sub-model to capture the features that are likely to exploit the bias and prevents the original model from learning these biased features by forcing orthogonality between these two sub-models. We performed evaluations on new balanced datasets extracted from the original MNLI dataset as well as the NLI stress tests, and show that the orthogonality approach is better at debiasing the model while maintaining competitive overall accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Generating Data to Mitigate Spurious Correlations in Natural Language Inference DatasetsYuxiang Wu, Matt Gardner, Pontus Stenetorp, Pradeep DasigiACL 2022 · 74 citations
- Looking at the Overlooked: An Analysis on the Word-Overlap Bias in Natural Language InferenceSara Rajaee, Yadollah Yaghoobzadeh, Mohammad Taher PilehvarEMNLP 2022 · 8 citations
- Improving the robustness of NLI models with minimax trainingMichalis Korakakis, Andreas VlachosACL 2023 · 4 citations
- Flexible Generation of Natural Language DeductionsKaj Bostrom, Xinyu Zhao, Swarat Chaudhuri, Greg DurrettEMNLP 2021 · 1 citation
- Resolving Lexical Bias in Model EditingHammad Rizwan, Domenic Rosati, Ga Wu, Hassan SajjadICML 2025
Builds on1
Related papers
- End-to-End Bias Mitigation by Modelling Biases in CorporaRabeeh Karimi Mahabadi, Yonatan Belinkov, James HendersonACL 2020 · 136 citations
- Towards Stable Natural Language Understanding via Information Entropy Guided DebiasingLi Du, Xiao Ding, Zhouhao Sun, Ting Liu et al.ACL 2023 · 1 citation
- Towards Debiasing NLU Models from Unknown BiasesPrasetya Ajie Utama, Nafise Sadat Moosavi, Iryna GurevychEMNLP 2020 · 3 citations
- Towards Understanding Gender Bias in Relation ExtractionAndrew Gaut, Tony Sun, Shirlyn Tang, Yuxin Huang et al.ACL 2020 · 7 citations
- A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector SpacesAnne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Ivan VulicAAAI 2020 · 68 citations
