"I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset
Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, Adina Williams
摘要
As language models grow in popularity, it becomes increasingly important to clearly measure all possible markers of demographic identity in order to avoid perpetuating existing societal harms. Many datasets for measuring bias currently exist, but they are restricted in their coverage of demographic axes and are commonly used with preset bias tests that presuppose which types of biases models can exhibit. In this work, we present a new, more inclusive bias measurement dataset, HOLIS-TICBIAS, which includes nearly 600 descriptor terms across 13 different demographic axes. HOLISTICBIAS was assembled in a participatory process including experts and community members with lived experience of these terms. These descriptors combine with a set of bias measurement templates to produce over 450,000 unique sentence prompts, which we use to explore, identify, and reduce novel forms of bias in several generative models. We demonstrate that HOLISTICBIAS is effective at measuring previously undetectable biases in token likelihoods from language models, as well as in an offensiveness classifier. We will invite additions and amendments to the dataset, which we hope will serve as a basis for more easy-to-use and standardized methods for evaluating bias in NLP models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Large Language Models are Geographically BiasedRohin Manvi, Samar Khanna, Marshall Burke, David B. Lobell 等ICML 2024 · 被引用 107 次
- Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language ModelsMyra Cheng, Esin Durmus, Dan JurafskyACL 2023 · 被引用 89 次
- WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language ModelsVirginia K. Felkner, Ho-Chun Herbert Chang, Eugene Jang, Jonathan MayACL 2023 · 被引用 46 次
- An Empathy-Based Sandbox Approach to Bridge the Privacy Gap among Attitudes, Goals, Knowledge, and BehaviorsChaoran Chen, Weijun Li, Wenxin Song, Yanfang Ye 等CHI 2024 · 被引用 25 次
- Measuring Political Bias in Large Language Models: What Is Said and How It Is SaidYejin Bang, Delong Chen, Nayeon Lee, Pascale FungACL 2024 · 被引用 21 次
它引用的顶会 Paper14
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
- Beyond Goldfish Memory: Long-Term Open-Domain ConversationJing Xu, Arthur Szlam, Jason WestonACL 2022 · 被引用 329 次
- Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language ModelsHannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal 等NeurIPS 2021 · 被引用 243 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
相关 Paper
- ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsDavid Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung 等EMNLP 2023 · 被引用 16 次
- Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language ModelsZara Siddique, Liam D. Turner, Luis Espinosa AnkeEMNLP 2024 · 被引用 2 次
- HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO DebiasingRuyi Chen, Lu Zhou, Xiaogang Xu, Chiyu Zhang 等ICML 2026 · 被引用 1 次
- Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in Languages at ScaleMarta R. Costa-jussà, Pierre Andrews, Eric Michael Smith, Prangthip Hansanti 等EMNLP 2023 · 被引用 2 次
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
