MISGENDERED: Limits of Large Language Models in Understanding Pronouns
Tamanna Hossain, Sunipa Dev, Sameer Singh
Abstract
Content Warning: This paper contains examples of misgendering and erasure that could be offensive and potentially triggering. Gender bias in language technologies has been widely studied, but research has mostly been restricted to a binary paradigm of gender. It is essential also to consider non-binary gender identities, as excluding them can cause further harm to an already marginalized group. In this paper, we comprehensively evaluate popular language models for their ability to correctly use English gender-neutral pronouns (e.g., singular they, them) and neo-pronouns (e.g., ze, xe, thon) that are used by individuals whose gender identity is not represented by binary pronouns. We introduce MISGENDERED, a framework for evaluating large language models' ability to correctly use preferred pronouns, consisting of (i) instances declaring an individual's pronoun, followed by a sentence with a missing pronoun, and (ii) an experimental setup for evaluating masked and auto-regressive language models using a unified method. When prompted outof-the-box, language models perform poorly at correctly predicting neo-pronouns (averaging 7.6% accuracy) and gender-neutral pronouns (averaging 31.0% accuracy). This inability to generalize results from a lack of representation of non-binary pronouns in training data and memorized associations. Few-shot adaptation with explicit examples in the prompt improves the performance but plateaus at only 45.4% for neo-pronouns. We release the full dataset, code, and demo at https://tamannahossainkay. github.io/misgendered/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db354546-6cd5-46d3-970d-dbc2635d27f8Cited by top-tier papers7
- Amplifying Trans and Nonbinary Voices: A Community-Centred Harm Taxonomy for LLMsEddie L. Ungless, Sunipa Dev, Cynthia L. Bennett, Rebecca Gulotta et al.ACL 2025 · 3 citations
- Gender Inclusivity Fairness Index (GIFI): A Multilevel Framework for Evaluating Gender Diversity in Large Language ModelsZhengyang Shan, Emily Diana, Jiawei ZhouACL 2025 · 3 citations
- Learning from Natural Language Explanations for Generalizable Entity MatchingSomin Wadhwa, Adit Krishnan, Runhui Wang, Byron C. Wallace et al.EMNLP 2024 · 3 citations
- LLM Bias Detection and Mitigation through the Lens of Desired DistributionsIngroj Shrestha, Padmini SrinivasanEMNLP 2025 · 3 citations
- The Lou Dataset - Exploring the Impact of Gender-Fair Language in German Text ClassificationAndreas Waldis, Joel Birrer, Anne Lauscher, Iryna GurevychEMNLP 2024 · 1 citation
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language ModelsHannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal et al.NeurIPS 2021 · 243 citations
- Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language TechnologiesSunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian et al.EMNLP 2021 · 113 citations
Related papers
- What about "em"? How Commercial Machine Translation Fails to Handle (Neo-)PronounsAnne Lauscher, Debora Nozza, Ehm Miltersen, Archie Crowley et al.ACL 2023 · 7 citations
- A Multilingual, Culture-First Approach to Addressing Misgendering in LLM ApplicationsSunayana Sitaram, Adrian de Wynter, Isobel McCrum, Qilong Gu et al.EMNLP 2025
- Are Models Biased on Text without Gender-related Language?Catarina G. Belém, Preethi Seshadri, Yasaman Razeghi, Sameer SinghICLR 2024 · 16 citations
- Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTEBeatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou et al.EMNLP 2025
- GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language ModelsKunsheng Tang, Wenbo Zhou, Jie Zhang, Aishan Liu et al.CCS 2024 · 7 citations
