"Fifty Shades of Bias": Normative Ratings of Gender Bias in GPT Generated English Text
Rishav Hada, Agrima Seth, Harshita Diddee, Kalika Bali
摘要
Warning: This paper contains statements that may be offensive or upsetting. Language serves as a powerful tool for the manifestation of societal belief systems. In doing so, it also perpetuates the prevalent biases in our society. Gender bias is one of the most pervasive biases in our society and is seen in online and offline discourses. With LLMs increasingly gaining human-like fluency in text generation, gaining a nuanced understanding of the biases these systems can generate is imperative. Prior work often treats gender bias as a binary classification task. However, acknowledging that bias must be perceived at a relative scale; we investigate the generation and consequent receptivity of manual annotators to bias of varying degrees. Specifically, we create the first dataset of GPT-generated English text with normative ratings of gender bias. Ratings were obtained using Best-Worst Scalingan efficient comparative annotation framework. Next, we systematically analyze the variation of themes of gender biases in the observed ranking and show that identity-attack is most closely related to gender bias. Finally, we show the performance of existing automated models trained on related concepts on our dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLPPieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett 等EMNLP 2024 · 被引用 4 次
- Attention Pruning: Automated Fairness Repair of Language Models via Surrogate Simulated AnnealingVishnu Asutosh Dasu, Md Rafi Ur Rashid, Vipul Gupta, Saeid Tizpaz-Niari 等ICSE 2026
- Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark DatasetsMahdi Zakizadeh, Mohammad Taher PilehvarEMNLP 2025
它引用的顶会 Paper7
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- Quantifying Intimacy in LanguageJiaxin Pei, David JurgensEMNLP 2020 · 被引用 44 次
- Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language ModelsHritik Bansal, John Dang, Aditya GroverICLR 2024 · 被引用 28 次
- Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark DatasetsSu Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim 等ACL 2021
相关 Paper
- Ruddit: Norms of Offensiveness for English Reddit CommentsRishav Hada, Sohi Sudhir, Pushkar Mishra, Helen Yannakoudakis 等ACL 2021
- Humans or LLMs as the Judge? A Study on Judgement BiasGuiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang 等EMNLP 2024 · 被引用 37 次
- Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion AttributionFlor Miriam Plaza del Arco, Amanda Cercas Curry, Alba Cercas Curry, Gavin Abercrombie 等ACL 2024 · 被引用 8 次
- Multi-Dimensional Gender Bias ClassificationEmily Dinan, Angela Fan, Ledell Wu, Jason Weston 等EMNLP 2020 · 被引用 7 次
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
