Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection
Suchin Gururangan, Dallas Card, Sarah K. Dreier, Emily K. Gade, Leroy Z. Wang, Zeyu Wang, Luke Zettlemoyer, Noah A. Smith
摘要
Language models increasingly rely on massive web crawls for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, and news often serve as anchors for automatically selecting web text most suitable for language modeling, a process typically referred to as quality filtering. Using a new dataset of U.S. high school newspaper articles-written by students from across the country-we investigate whose language is preferred by the quality filter used for GPT-3. We find that newspapers from larger schools, located in wealthier, educated, and urban zones (ZIP codes) are more likely to be classified as high quality. We also show that this quality measurement is unaligned with other sensible metrics, such as factuality or literary acclaim. We argue that privileging any corpus as high quality entails a language ideology, and more care is needed to construct training corpora for language models, with better transparency and justification for the inclusion or exclusion of various texts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 被引用 301 次
- Prompting PaLM for Translation: Assessing Strategies and PerformanceDavid Vilar, Markus Freitag, Colin Cherry, Jiaming Luo 等ACL 2023 · 被引用 70 次
- Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4Kent K. Chang, Mackenzie Cramer, Sandeep Soni, David BammanEMNLP 2023 · 被引用 70 次
- Do All Languages Cost the Same? Tokenization in the Era of Commercial Language ModelsOrevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai 等EMNLP 2023 · 被引用 24 次
- NLPositionality: Characterizing Design Biases of Datasets and ModelsSebastin Santy, Jenny T. Liang, Ronan Le Bras, Katharina Reinecke 等ACL 2023 · 被引用 23 次
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- HTLM: Hyper-Text Pre-Training and Prompting of Language ModelsArmen Aghajanyan, Dmytro Okhonko, Mike Lewis, Mandar Joshi 等ICLR 2022 · 被引用 82 次
相关 Paper
- How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLPKushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger 等ACL 2026 · 被引用 2 次
- AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data FiltersLi Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell 等ACL 2024 · 被引用 2 次
- Removing Noise, not Finding Gold: Quality Filtering for Large-Scale PretrainingThiziri Nait Saada, Louis Béthune, Michal Klein, David Grangier 等ICML 2026
- QuRating: Selecting High-Quality Data for Training Language ModelsAlexander Wettig, Aatmik Gupta, Saumya Malik, Danqi ChenICML 2024 · 被引用 138 次
- The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite GraphMinghao Wu, Thuy-Trang Vu, Lizhen Qu, Gholamreza HaffariICML 2025
