Mitigating Spurious Correlation in Natural Language Understanding with Counterfactual Inference
Can Udomcharoenchaikit, Wuttikorn Ponwitayarat, Patomporn Payoungkhamdee, Kanruethai Masuk, Weerayut Buaphet, Ekapol Chuangsuwanich, Sarana Nutanong
Abstract
Despite their promising results on standard benchmarks, NLU models are still prone to make predictions based on shortcuts caused by unintended bias in the dataset. For example, an NLI model may use lexical overlap as a shortcut to make entailment predictions due to repetitive data generation patterns from annotators, also called annotation artifacts. In this paper, we propose a causal analysis framework to help debias NLU models. We show that (1) by defining causal relationships, we can introspect how much annotation artifacts affect the outcomes. (2) We can utilize counterfactual inference to mitigate bias with this knowledge. We found that viewing a model as a treatment can mitigate bias more effectively than viewing annotation artifacts as treatment. (3) In addition to bias mitigation, we can interpret how much each debiasing strategy is affected by annotation artifacts. Our experimental results show that using counterfactual inference can improve out-of-distribution performance in all settings while maintaining high in-distribution performance. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 413fb1dd-c37c-449d-bcf5-32c59f5b4745Cited by top-tier papers4
- Improving Bias Mitigation through Bias Experts in Natural Language UnderstandingEojin Jeon, Mingyu Lee, Juhyeong Park, Yeachan Kim et al.EMNLP 2023 · 3 citations
- Focus On This, Not That! Steering LLMs with Adaptive Feature SpecificationTom A. Lamb, Adam Davies, Alasdair Paren, Philip Torr et al.ICML 2025
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic TriplesKyohoon Jin, Juhwan Choi, Jungmin Yun, Junho Lee et al.EMNLP 2025
- Curriculum Debiasing: Toward Robust Parameter-Efficient Fine-Tuning Against Dataset BiasesMingyu Lee, Yeachan Kim, Wing-Lam Mok, SangKeun LeeACL 2025
Builds on14
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Model-Agnostic Counterfactual Reasoning for Eliminating Popularity Bias in Recommender SystemTianxin Wei, Fuli Feng, Jiawei Chen, Ziwei Wu et al.KDD 2021 · 246 citations
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers et al.ICML 2020 · 242 citations
- Clicks can be Cheating: Counterfactual Recommendation for Mitigating Clickbait IssueWenjie Wang, Fuli Feng, Xiangnan He, Hanwang Zhang et al.SIGIR 2021 · 173 citations
Related papers
- Debiasing NLU Models via Causal Intervention and Counterfactual ReasoningBing Tian, Yixin Cao, Yong Zhang, Chunxiao XingAAAI 2022 · 45 citations
- Counterfactual VQA: A Cause-Effect Look at Language BiasYulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu et al.CVPR 2021
- Towards Debiasing NLU Models from Unknown BiasesPrasetya Ajie Utama, Nafise Sadat Moosavi, Iryna GurevychEMNLP 2020 · 3 citations
- Supervising Model Attention with Human Explanations for Robust Natural Language InferenceJoe Stacey, Yonatan Belinkov, Marek ReiAAAI 2022 · 52 citations
- Interventional Training for Out-Of-Distribution Natural Language UnderstandingSicheng Yu, Jing Jiang, Hao Zhang, Yulei Niu et al.EMNLP 2022 · 3 citations
