AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren F. Klein, Jesse Dodge
Abstract
Large language models' (LLMs) abilities are drawn from their pretraining data, and model development begins with data curation. However, decisions around what data is retained or removed during this initial stage are underscrutinized. In our work, we ground web text, which is a popular pretraining data source, to its social and geographic contexts. We create a new dataset of 10.3 million self-descriptions of website creators, and extract information about who they are and where they are from: their topical interests, social roles, and geographic affiliations. Then, we conduct the first study investigating how ten "quality" and English language identification (langID) filters affect webpages that vary along these social dimensions. Our experiments illuminate a range of implicit preferences in data curation: we show that some quality classifiers act like topical domain filters, and langID can overlook English content from some regions of the world. Overall, we hope that our work will encourage a new line of research on pretraining data curation practices and its social implications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4aeeba5d-e87d-4b10-8dbe-4dcc0309d019Cited by top-tier papers8
- Kaleidoscope: In-language Exams for Massively Multilingual Vision EvaluationIsrafel Salazar, Manuel Fernández Burda, Shayekh Bin Islam, Arshia Soltani Moakhar et al.ICLR 2026 · 8 citations
- KidLM: Advancing Language Models for Children - Early Insights and Future DirectionsMir Tafseer Nayeem, Davood RafieiEMNLP 2024 · 7 citations
- Enhancing LLMs via High-Knowledge Data SelectionFeiyu Duan, Xuemiao Zhang, Sirui Wang, Haoran Que et al.AAAI 2025 · 3 citations
- so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMsSriharsh Bhyravajjula, Melanie Walsh, Anna Preus, Maria AntoniakEMNLP 2025 · 1 citation
- Aya Dataset: An Open-Access Collection for Multilingual Instruction TuningShivalika Singh, Freddie Vargus, Daniel D'souza, Börje Karlsson et al.ACL 2024
Builds on8
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace et al.ICML 2023 · 623 citations
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 383 citations
- Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language TechnologiesSunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian et al.EMNLP 2021 · 113 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
Related papers
- Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data SelectionSuchin Gururangan, Dallas Card, Sarah K. Dreier, Emily K. Gade et al.EMNLP 2022 · 41 citations
- What Makes a High-Quality Training Dataset for Large Language Models: A Practitioners' PerspectiveXiao Yu, Zexian Zhang, Feifei Niu, Xing Hu et al.ASE 2024 · 15 citations
- Data, Data Everywhere: A Guide for Pretraining Dataset ConstructionJupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Bo Liu et al.EMNLP 2024
- Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining DataHaoran Deng, Yingyu Lin, Zhenghao Lin, Xiao Liu et al.ICLR 2026 · 6 citations
- Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language ModelsMehdi Ali, Manuel Brack, Max Lübbering, Elias Wendt et al.EMNLP 2025 · 1 citation
