FairPrism: Evaluating Fairness-Related Harms in Text Generation
Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, Hanna M. Wallach
摘要
It is critical to measure and mitigate fairnessrelated harms caused by AI text generation systems, including stereotyping and demeaning harms. To that end, we introduce FairPrism, a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harms relating to gender and sexuality. FairPrism aims to address several limitations of existing datasets for measuring and mitigating fairness-related harms, including improved transparency, clearer specification of dataset coverage, and accounting for annotator disagreement and harms that are context-dependent. FairPrism's annotations include the extent of stereotyping and demeaning harms, the demographic groups targeted, and appropriateness for different applications. The annotations also include specific harms that occur in interactive contexts and harms that raise normative concerns when the "speaker" is an AI system. Due to its precision and granularity, FairPrism can be used to diagnose (1) the types of fairnessrelated harms that AI text generation systems cause, and (2) the potential limitations of mitigation methods, both of which we illustrate through case studies. Finally, the process we followed to develop FairPrism offers a recipe for building improved datasets for measuring and mitigating harms caused by AI systems. Real Toxicity Prompts BOLD ToxiGen Social Bias Frames Our work: FairPrism Text source AI AI AI Human AI Label source (human or classifier) Classifier Classifier Classifier (792 human) Human Human Separates subtypes of harm within toxicity/hate speech? (3.1) No No No No Yes Contextualizes AI responses? (3) No Yes No N/A Yes Identifies target group harmed? (3.2) No No Yes Yes Yes Includes disaggregated data? (3.3) No No No Yes Yes Examines AI-specific harms? (3.4) No No No No Yes
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language ModelsYan Liu, Yu Liu, Xiaokang Chen, Pin-Yu Chen 等ICLR 2024 · 被引用 32 次
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 被引用 11 次
- Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a BudgetFlorian E. Dorner, Moritz HardtICML 2024 · 被引用 10 次
- GuardBench: A Large-Scale Benchmark for Guardrail ModelsElias Bassani, Ignacio SanchezEMNLP 2024 · 被引用 7 次
- ECBD: Evidence-Centered Benchmark Design for NLPYu Lu Liu, Su Lin Blodgett, Jackie C. K. Cheung, Vera Liao 等ACL 2024 · 被引用 3 次
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi 等EMNLP 2021 · 被引用 159 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Detection in Conversational AIAmanda Cercas Curry, Gavin Abercrombie, Verena RieserEMNLP 2021 · 被引用 38 次
- When Are Search Completion Suggestions Problematic?Alexandra Olteanu, Fernando Diaz, Gabriella KazaiCSCW 2020 · 被引用 38 次
相关 Paper
- AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness BenchmarkLi Lin, Santosh Santosh, Mingyang Wu, Xin Wang 等CVPR 2025
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 被引用 495 次
- Multi-Dimensional Gender Bias ClassificationEmily Dinan, Angela Fan, Ledell Wu, Jason Weston 等EMNLP 2020 · 被引用 7 次
- ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsDavid Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung 等EMNLP 2023 · 被引用 16 次
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
