FairPrism: Evaluating Fairness-Related Harms in Text Generation
Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, Hanna M. Wallach
Abstract
It is critical to measure and mitigate fairnessrelated harms caused by AI text generation systems, including stereotyping and demeaning harms. To that end, we introduce FairPrism, a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harms relating to gender and sexuality. FairPrism aims to address several limitations of existing datasets for measuring and mitigating fairness-related harms, including improved transparency, clearer specification of dataset coverage, and accounting for annotator disagreement and harms that are context-dependent. FairPrism's annotations include the extent of stereotyping and demeaning harms, the demographic groups targeted, and appropriateness for different applications. The annotations also include specific harms that occur in interactive contexts and harms that raise normative concerns when the "speaker" is an AI system. Due to its precision and granularity, FairPrism can be used to diagnose (1) the types of fairnessrelated harms that AI text generation systems cause, and (2) the potential limitations of mitigation methods, both of which we illustrate through case studies. Finally, the process we followed to develop FairPrism offers a recipe for building improved datasets for measuring and mitigating harms caused by AI systems. Real Toxicity Prompts BOLD ToxiGen Social Bias Frames Our work: FairPrism Text source AI AI AI Human AI Label source (human or classifier) Classifier Classifier Classifier (792 human) Human Human Separates subtypes of harm within toxicity/hate speech? (3.1) No No No No Yes Contextualizes AI responses? (3) No Yes No N/A Yes Identifies target group harmed? (3.2) No No Yes Yes Yes Includes disaggregated data? (3.3) No No No Yes Yes Examines AI-specific harms? (3.4) No No No No Yes
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ada9b7ac-0136-4e15-848f-3d81cdf69bc9Cited by top-tier papers7
- The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language ModelsYan Liu, Yu Liu, Xiaokang Chen, Pin-Yu Chen et al.ICLR 2024 · 32 citations
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 11 citations
- Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a BudgetFlorian E. Dorner, Moritz HardtICML 2024 · 10 citations
- GuardBench: A Large-Scale Benchmark for Guardrail ModelsElias Bassani, Ignacio SanchezEMNLP 2024 · 7 citations
- ECBD: Evidence-Centered Benchmark Design for NLPYu Lu Liu, Su Lin Blodgett, Jackie C. K. Cheung, Vera Liao et al.ACL 2024 · 3 citations
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi et al.EMNLP 2021 · 159 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Detection in Conversational AIAmanda Cercas Curry, Gavin Abercrombie, Verena RieserEMNLP 2021 · 38 citations
- When Are Search Completion Suggestions Problematic?Alexandra Olteanu, Fernando Diaz, Gabriella KazaiCSCW 2020 · 38 citations
Related papers
- AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness BenchmarkLi Lin, Santosh Santosh, Mingyang Wu, Xin Wang et al.CVPR 2025
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 495 citations
- Multi-Dimensional Gender Bias ClassificationEmily Dinan, Angela Fan, Ledell Wu, Jason Weston et al.EMNLP 2020 · 7 citations
- ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsDavid Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung et al.EMNLP 2023 · 16 citations
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
