Data Caricatures: On the Representation of African American Language in Pretraining Corpora
Nicholas Deas, Blake Vente, Amith Ananthram, Jessica Grieser, Desmond Upton Patton, Shana Kleiner, James R. Shepard III, Kathleen McKeown
摘要
With a combination of quantitative experiments, human judgments, and qualitative analyses, we evaluate the quantity and quality of African American Language (AAL) representation in 12 predominantly English, open-source pretraining corpora. We specifically focus on the sources, variation, and naturalness of included AAL texts representing the AALspeaking community. We find that AAL is underrepresented in all evaluated pretraining corpora compared to US demographics, constituting as few as 0.007% and at most 0.18% of documents. We also find that more than 25% of AAL texts in C4 may be perceived as inappropriate for LLMs to generate and to reinforce harmful stereotypes. Finally, we find that most automated filters are more likely to conserve White Mainstream English (WME) texts over AAL in pretraining corpora. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- QuRating: Selecting High-Quality Data for Training Language ModelsAlexander Wettig, Aatmik Gupta, Saumya Malik, Danqi ChenICML 2024 · 被引用 138 次
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 被引用 117 次
- How to train data-efficient LLMsNoveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang, Jianmo Ni 等ICLR 2026 · 被引用 106 次
- "It's Kind of Like Code-Switching": Black Older Adults' Experiences with a Voice Assistant for Health Information SeekingChristina N. Harrington, Radhika Garg, Amanda T. Woodward, Dimitri WilliamsCHI 2022 · 被引用 99 次
相关 Paper
- Evaluation of African American Language Bias in Natural Language GenerationNicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton 等EMNLP 2023 · 被引用 15 次
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentCameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim 等ICML 2026 · 被引用 22 次
- Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for ChatbotsSarah E. Finch, Ellie S. Paek, Ikseon Choi, Jinho D. ChoiACL 2025
- Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language ModelsRobert Wolfe, Aylin CaliskanEMNLP 2021 · 被引用 27 次
- AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African LanguagesHao Yu, Tianyi Xu, Michael A. Hedderich, Wassim Hamidouche 等ACL 2026 · 被引用 5 次
