Predictive Biases in Natural Language Processing Models: A Conceptual Framework and Overview
Deven Shah, H. Andrew Schwartz, Dirk Hovy
Abstract
An increasing number of natural language processing papers address the effect of bias on predictions, introducing mitigation techniques at different parts of the standard NLP pipeline (data and models). However, these works have been conducted individually, without a unifying framework to organize efforts within the field. This situation leads to repetitive approaches, and focuses overly on bias symptoms/effects, rather than on their origins, which could limit the development of effective countermeasures. In this paper, we propose a unifying predictive bias framework for NLP. We summarize the NLP literature and suggest general mathematical definitions of predictive bias. We differentiate two consequences of bias: outcome disparities and error disparities, as well as four potential origins of biases: label bias, selection bias, model overamplification, and semantic bias. Our framework serves as an overview of predictive bias in NLP, integrating existing work into a single structure, and providing a conceptual baseline for improved frameworks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eac642a1-3dc2-4cb3-b28d-ab4258377c79Cited by top-tier papers29
- Competency Problems: On Finding and Removing Artifacts in Language DataMatt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters et al.EMNLP 2021 · 72 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- Debiasing NLU Models via Causal Intervention and Counterfactual ReasoningBing Tian, Yixin Cao, Yong Zhang, Chunxiao XingAAAI 2022 · 45 citations
- Why So Toxic?: Measuring and Triggering Toxic Behavior in Open-Domain ChatbotsWai Man Si, Michael Backes, Jeremy Blackburn, Emiliano De Cristofaro et al.CCS 2022 · 34 citations
- Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech TranslationBeatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri et al.ACL 2022 · 30 citations
Builds on1
Related papers
- Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLPPieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett et al.EMNLP 2024 · 4 citations
- Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant LearningFan Zhou, Yuzhou Mao, Liu Yu, Yi Yang et al.ACL 2023 · 21 citations
- A Survey of Race, Racism, and Anti-Racism in NLPAnjalie Field, Su Lin Blodgett, Zeerak Waseem, Yulia TsvetkovACL 2021
- DPA: A one-stop metric to measure bias amplification in classification datasetsBhanu Tokas, Rahul Nair, Hannah KernerNeurIPS 2025 · 1 citation
- Quantifying Societal Bias Amplification in Image CaptioningYusuke Hirota, Yuta Nakashima, Noa GarciaCVPR 2022 · 45 citations
