Evaluation of African American Language Bias in Natural Language Generation
Nicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton, Elsbeth Turcan, Kathleen R. McKeown
Abstract
Warning: This paper contains content and language that may be considered offensive to some readers. While biases disadvantaging African American Language (AAL) have been uncovered in models for tasks such as speech recognition and toxicity detection, there has been little investigation of these biases for language generation models like ChatGPT. We evaluate how well LLMs understand AAL in comparison to White Mainstream English (WME), the encouraged "standard" form of English taught in American classrooms. We measure large language model performance on two tasks: a counterpart generation task, where a model generates AAL given WME and vice versa, as well as a masked span prediction (MSP) task, where models predict a phrase hidden from their input. Using a novel dataset of AAL texts from a variety of regions and contexts, we present evidence of dialectal bias for six pre-trained LLMs through performance gaps on these tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1642daa8-1490-43ed-9225-d317c8d5d717Cited by top-tier papers10
- Linguistic Bias in ChatGPT: Language Models Reinforce Dialect DiscriminationEve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi et al.EMNLP 2024 · 36 citations
- Guiding LLM Decision-Making with Fairness Reward ModelsZara Hall, Melanie Subbiah, Thomas P. Zollo, Kathleen McKeown et al.NeurIPS 2025 · 14 citations
- Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning TasksFangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann et al.ACL 2025 · 14 citations
- Should AI Mimic People? Understanding AI-Supported Writing Technology Among Black UsersJeffrey Basoah, Jay L. Cunningham, Erica Adams, Alisha Bose et al.CSCW 2025 · 5 citations
- ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in ContextVictoria R. Li, Yida Chen, Naomi SaphraEMNLP 2024 · 5 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Predictive Biases in Natural Language Processing Models: A Conceptual Framework and OverviewDeven Shah, H. Andrew Schwartz, Dirk HovyACL 2020 · 93 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
Related papers
- Data Caricatures: On the Representation of African American Language in Pretraining CorporaNicholas Deas, Blake Vente, Amith Ananthram, Jessica Grieser et al.ACL 2025
- Are Stereotypes Leading LLMs' Zero-Shot Stance Detection ?Anthony Dubreuil, Antoine Gourru, Christine Largeron, Amine TrabelsiEMNLP 2025 · 1 citation
- MEGA: Multilingual Evaluation of Generative AIKabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng et al.EMNLP 2023 · 91 citations
- Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for ChatbotsSarah E. Finch, Ellie S. Paek, Ikseon Choi, Jinho D. ChoiACL 2025
- Identifying, Explaining, and Correcting Ableist Language with AIKynnedy Simone Smith, Lydia B. Chilton, Danielle BraggCHI 2026 · 1 citation
