Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLP
Pieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett, Zeerak Talat
Abstract
This paper introduces the concept of actionability in the context of bias measures in natural language processing (NLP). We define actionability as the degree to which a measurement's results enable informed action and propose a set of desiderata for assessing it. Building on existing frameworks such as measurement modeling, we argue that actionability is a crucial aspect of bias measures that has been largely overlooked in the literature. We conduct a comprehensive review of 146 papers proposing bias measures in NLP, examining whether and how they provide the information required for actionable results. Our findings reveal that many key elements of actionability, including a measure's intended use and reliability assessment, are often unclear or absent. This study highlights a significant gap in the current approach to developing and reporting bias measures in NLP. We argue that this lack of clarity may impede the effective implementation and utilization of these measures. To address this issue, we offer recommendations for more comprehensive and actionable metric development and reporting practices in NLP bias research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae122a8c-43b3-4fa0-aa4b-61698bc3eab2Cited by top-tier papers3
- Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party EvaluationsAnka Reuel, Avijit Ghosh, Jenny Chim, Andrew Tran et al.ICML 2026 · 9 citations
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising OpinionsNannan Huang, Haytham M. Fayek, Xiuzhen ZhangEMNLP 2025 · 2 citations
- PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training RunsOskar van der Wal, Pietro Lesci, Max Müller-Eberstein, Naomi Saphra et al.ICLR 2025
Builds on46
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than EnglishAurélie Névéol, Yoann Dupont, Julien Bezançon, Karën FortACL 2022 · 61 citations
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani et al.EMNLP 2022 · 56 citations
- Perturbation Augmentation for Fairer NLPRebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith et al.EMNLP 2022 · 54 citations
- CoMPosT: Characterizing and Evaluating Caricature in LLM SimulationsMyra Cheng, Tiziano Piccardi, Diyi YangEMNLP 2023 · 33 citations
Related papers
- Predictive Biases in Natural Language Processing Models: A Conceptual Framework and OverviewDeven Shah, H. Andrew Schwartz, Dirk HovyACL 2020 · 93 citations
- Modeling Disclosive Transparency in NLP Application DescriptionsMichael Saxon, Sharon Levy, Xinyi Wang, Alon Albalak et al.EMNLP 2021 · 3 citations
- Intrinsic Bias Metrics Do Not Correlate with Application BiasSeraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya et al.ACL 2021
- A Survey of Race, Racism, and Anti-Racism in NLPAnjalie Field, Su Lin Blodgett, Zeerak Waseem, Yulia TsvetkovACL 2021
- Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement TheoryZiang Xiao, Susu Zhang, Vivian Lai, Q. Vera LiaoEMNLP 2023 · 6 citations
