AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren F. Klein, Jesse Dodge
摘要
Large language models' (LLMs) abilities are drawn from their pretraining data, and model development begins with data curation. However, decisions around what data is retained or removed during this initial stage are underscrutinized. In our work, we ground web text, which is a popular pretraining data source, to its social and geographic contexts. We create a new dataset of 10.3 million self-descriptions of website creators, and extract information about who they are and where they are from: their topical interests, social roles, and geographic affiliations. Then, we conduct the first study investigating how ten "quality" and English language identification (langID) filters affect webpages that vary along these social dimensions. Our experiments illuminate a range of implicit preferences in data curation: we show that some quality classifiers act like topical domain filters, and langID can overlook English content from some regions of the world. Overall, we hope that our work will encourage a new line of research on pretraining data curation practices and its social implications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Kaleidoscope: In-language Exams for Massively Multilingual Vision EvaluationIsrafel Salazar, Manuel Fernández Burda, Shayekh Bin Islam, Arshia Soltani Moakhar 等ICLR 2026 · 被引用 8 次
- KidLM: Advancing Language Models for Children - Early Insights and Future DirectionsMir Tafseer Nayeem, Davood RafieiEMNLP 2024 · 被引用 7 次
- Enhancing LLMs via High-Knowledge Data SelectionFeiyu Duan, Xuemiao Zhang, Sirui Wang, Haoran Que 等AAAI 2025 · 被引用 3 次
- so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMsSriharsh Bhyravajjula, Melanie Walsh, Anna Preus, Maria AntoniakEMNLP 2025 · 被引用 1 次
- Aya Dataset: An Open-Access Collection for Multilingual Instruction TuningShivalika Singh, Freddie Vargus, Daniel D'souza, Börje Karlsson 等ACL 2024
它引用的顶会 Paper8
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace 等ICML 2023 · 被引用 623 次
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 被引用 383 次
- Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language TechnologiesSunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian 等EMNLP 2021 · 被引用 113 次
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali 等ACL 2020 · 被引用 40 次
相关 Paper
- Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data SelectionSuchin Gururangan, Dallas Card, Sarah K. Dreier, Emily K. Gade 等EMNLP 2022 · 被引用 41 次
- What Makes a High-Quality Training Dataset for Large Language Models: A Practitioners' PerspectiveXiao Yu, Zexian Zhang, Feifei Niu, Xing Hu 等ASE 2024 · 被引用 15 次
- Data, Data Everywhere: A Guide for Pretraining Dataset ConstructionJupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Bo Liu 等EMNLP 2024
- Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining DataHaoran Deng, Yingyu Lin, Zhenghao Lin, Xiao Liu 等ICLR 2026 · 被引用 6 次
- Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language ModelsMehdi Ali, Manuel Brack, Max Lübbering, Elias Wendt 等EMNLP 2025 · 被引用 1 次
