Debiasing Methods in Natural Language Understanding Make Bias More Accessible
Michael Mendelson, Yonatan Belinkov
摘要
Model robustness to bias is often determined by the generalization on carefully designed out-of-distribution datasets. Recent debiasing methods in natural language understanding (NLU) improve performance on such datasets by pressuring models into making unbiased predictions. An underlying assumption behind such methods is that this also leads to the discovery of more robust features in the model's inner representations. We propose a general probing-based framework that allows for posthoc interpretation of biases in language models, and use an information-theoretic approach to measure the extractability of certain biases from the model's representations. We experiment with several NLU datasets and known biases, and show that, counter-intuitively, the more a language model is pushed towards a debiased regime, the more bias is actually encoded in its inner representations. 1 * Supported by the Viterbi Fellowship in the Center for Computer Engineering at the Technion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for FreeHao Li, Xiaogeng Liu, Ning Zhang, Chaowei XiaoACL 2025 · 被引用 33 次
- Are All Spurious Features in Natural Language Alike? An Analysis through a Causal LensNitish Joshi, Xiang Pan, He HeEMNLP 2022 · 被引用 19 次
- Feature-Level Debiased Natural Language UnderstandingYougang Lyu, Piji Li, Yechang Yang, Maarten de Rijke 等AAAI 2023 · 被引用 12 次
- Looking at the Overlooked: An Analysis on the Word-Overlap Bias in Natural Language InferenceSara Rajaee, Yadollah Yaghoobzadeh, Mohammad Taher PilehvarEMNLP 2022 · 被引用 8 次
- BLIND: Bias Removal With No DemographicsHadas Orgad, Yonatan BelinkovACL 2023 · 被引用 8 次
它引用的顶会 Paper5
- End-to-End Bias Mitigation by Modelling Biases in CorporaRabeeh Karimi Mahabadi, Yonatan Belinkov, James HendersonACL 2020 · 被引用 136 次
- Learning from others' mistakes: Avoiding dataset biases without modeling themVictor Sanh, Thomas Wolf, Yonatan Belinkov, Alexander M. RushICLR 2021 · 被引用 123 次
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 被引用 34 次
- Mind the Trade-off: Debiasing NLU Models without Degrading the In-distribution PerformancePrasetya Ajie Utama, Nafise Sadat Moosavi, Iryna GurevychACL 2020 · 被引用 11 次
- Towards Debiasing NLU Models from Unknown BiasesPrasetya Ajie Utama, Nafise Sadat Moosavi, Iryna GurevychEMNLP 2020 · 被引用 3 次
相关 Paper
- Towards Stable Natural Language Understanding via Information Entropy Guided DebiasingLi Du, Xiao Ding, Zhouhao Sun, Ting Liu 等ACL 2023 · 被引用 1 次
- Probing as Quantifying Inductive BiasAlexander Immer, Lucas Torroba Hennigen, Vincent Fortuin, Ryan CotterellACL 2022
- Predicting Fine-Tuning Performance with ProbingZining Zhu, Soroosh Shahtalebi, Frank RudziczEMNLP 2022 · 被引用 6 次
- Debiasing NLU Models via Causal Intervention and Counterfactual ReasoningBing Tian, Yixin Cao, Yong Zhang, Chunxiao XingAAAI 2022 · 被引用 45 次
- Monitoring Latent World States in Language Models with Propositional ProbesJiahai Feng, Stuart Russell, Jacob SteinhardtICLR 2025
