WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models
Virginia K. Felkner, Ho-Chun Herbert Chang, Eugene Jang, Jonathan May
Abstract
Content Warning: This paper contains examples of homophobic and transphobic stereotypes. We present WinoQueer: a benchmark specifically designed to measure whether large language models (LLMs) encode biases that are harmful to the LGBTQ+ community. The benchmark is community-sourced, via application of a novel method that generates a bias benchmark from a community survey. We apply our benchmark to several popular LLMs and find that off-the-shelf models generally do exhibit considerable anti-queer bias. Finally, we show that LLM bias against a marginalized community can be somewhat mitigated by finetuning on data written about or by members of that community, and that social media text written by community members is more effective than news text written about the community by non-members. Our method for community-in-the-loop benchmark development provides a blueprint for future researchers to develop community-driven, harms-grounded LLM benchmarks for other marginalized communities. Note: This version corrects a bug found in evaluation code after publication. General findings have not changed, but tables 5 and 6 and figure 1 have been corrected.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 124a862e-dc0a-4a15-9ead-b7e7525eb69aCited by top-tier papers18
- Design Principles for Generative AI ApplicationsJustin D. Weisz, Jessica He, Michael J. Muller, Gabriela Hoefer et al.CHI 2024 · 221 citations
- Evaluating the Experience of LGBTQ+ People Using Large Language Model Based Chatbots for Mental Health SupportZilin Ma, Yiyang Mei, Yinru Long, Zhaoyuan Su et al.CHI 2024 · 70 citations
- A Piece of Theatre: Investigating How Teachers Design LLM Chatbots to Assist Adolescent Cyberbullying EducationMichael A. Hedderich, Natalie N. Bazarova, Wenting Zou, Ryun Shim et al.CHI 2024 · 42 citations
- GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language ModelsTao Zhang, Ziqian Zeng, Yuxiang Xiao, Huiping Zhuang et al.ACL 2025 · 18 citations
- GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language ModelsKunsheng Tang, Wenbo Zhou, Jie Zhang, Aishan Liu et al.CCS 2024 · 7 citations
Builds on8
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than EnglishAurélie Névéol, Yoann Dupont, Julien Bezançon, Karën FortACL 2022 · 61 citations
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani et al.EMNLP 2022 · 56 citations
Related papers
- GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark ConstructionVirginia K. Felkner, Jennifer A. Thompson, Jonathan MayACL 2024 · 3 citations
- Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language ModelsZara Siddique, Liam D. Turner, Luis Espinosa AnkeEMNLP 2024 · 2 citations
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 495 citations
- Amplifying Trans and Nonbinary Voices: A Community-Centred Harm Taxonomy for LLMsEddie L. Ungless, Sunipa Dev, Cynthia L. Bennett, Rebecca Gulotta et al.ACL 2025 · 3 citations
- Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre-Trained Language ModelsRyan Steed, Swetasudha Panda, Ari Kobren, Michael L. WickACL 2022 · 52 citations
