Organize the Web: Constructing Domains Enhances Pre-Training Data Curation
Alexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi, Danqi Chen, Luca Soldaini
摘要
Modern language models are trained on large, unstructured datasets consisting of trillions of tokens and obtained by crawling the web. The unstructured nature makes it difficult to reason about their contents and develop systematic approaches to data curation. In this paper, we unpack monolithic web corpora by developing taxonomies of their contents and organizing them into domains. We introduce WebOrganizer, a framework for organizing web pages in terms of both their topic and format. Using these two complementary notions of domains, we automatically annotate pre-training data by distilling annotations from a large language model into efficient classifiers. This allows us to study how data from different domains should be mixed to improve models on downstream tasks, and we show that we can combine insights about effective topics and formats to further boost performance. We demonstrate that our domain mixing also improves existing methods that select data based on quality. Furthermore, we study and compare how quality-based methods will implicitly change the domain mixture. Overall, our work demonstrates that constructing and mixing domains provides a valuable complement to quality-based data curation methods, opening new avenues for effective and insightful pre-training data curation. Website weborganizer.allen.ai Models & Data huggingface.co/WebOrganizer Code github.com/CodeCreator/WebOrganizer
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang 等ICLR 2026 · 被引用 250 次
- OpenThoughts: Data Recipes for Reasoning ModelsEtash Kumar Guha, Ryan Marten, Sedrick Keh, Negin Raoof 等ICLR 2026 · 被引用 235 次
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model EvaluationDavid Heineman, Valentin Hofmann, Ian Magnusson, Yuling Gu 等NeurIPS 2025 · 被引用 24 次
- FlexOLMo: Open Language Models for Flexible Data UseWeijia Shi, Akshita Bhagia, Kevin Farhat, Niklas Muennighoff 等NeurIPS 2025 · 被引用 16 次
- Group-Level Data Selection for Efficient PretrainingZichun Yu, Fei Peng, Jie Lei, Arnold Overwijk 等NeurIPS 2025 · 被引用 13 次
它引用的顶会 Paper32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
相关 Paper
- Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without TrainingMozhi Zhang, Howe Tissue, Lu Wang, Xipeng QiuICML 2025
- Webformer: Pre-training with Web Pages for Information RetrievalYu Guo, Zhengyi Ma, Jiaxin Mao, Hongjin Qian 等SIGIR 2022 · 被引用 30 次
- RegMix: Data Mixture as Regression for Language Model Pre-trainingQian Liu, Xiaosen Zheng, Niklas Muennighoff, Guangtao Zeng 等ICLR 2025
- AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data FiltersLi Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell 等ACL 2024 · 被引用 2 次
- DataMan: Data Manager for Pre-training Large Language ModelsRu Peng, Kexin Yang, Yawen Zeng, Junyang Lin 等ICLR 2025
