Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
Fan Zhou, Zengzhi Wang, Qian Liu, Junlong Li, Pengfei Liu
摘要
Large language model pre-training has traditionally relied on human experts to craft heuristics for improving the corpora quality, resulting in numerous rules developed to date. However, these rules lack the flexibility to address the unique characteristics of individual example effectively. Meanwhile, applying tailored rules to every example is impractical for human experts. In this paper, we demonstrate that even small language models, with as few as 0.3B parameters, can exhibit substantial data refining capabilities comparable to those of human experts. We introduce Programming Every Example (PROX), a novel framework that treats data refinement as a programming task, enabling models to refine corpora by generating and executing fine-grained operations, such as string normalization, for each individual example at scale. Experimental results show that models pre-trained on PROX-curated data outperform either original data or data filtered by other selection methods by more than 2% across various downstream benchmarks. Its effectiveness spans various model sizes and pre-training corpora, including C4, RedPajama-V2, FineWeb, FineWeb-Edu, and DCLM. Furthermore, PROX exhibits significant potential in domain-specific continual pre-training: without domain specific design, models trained on OpenWebMath refined by PROX outperform human-crafted rule-based methods, improving average accuracy by 7.6% over MISTRAL-7B, with 14.6% for LLAMA-2-7B and 20.3% for CODELLAMA-7B, all within 10B tokens to be comparable to models like LLEMMA-7B trained on 200B tokens. Further analysis highlights that PROX significantly saves training FLOPs, offering a promising path for efficient LLM pre-training. We are open-sourcing PROX with ≥ 500B corpus, models, and sharing all training and implementation details for reproducible research and future innovation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer 等NeurIPS 2025 · 被引用 431 次
- DiffuCoder: Understanding and Improving Masked Diffusion Models for Code GenerationShansan Gong, Ruixiang Zhang, Huangjie Zheng, Jiatao Gu 等ICLR 2026 · 被引用 198 次
- Tool Learning in the Wild: Empowering Language Models as Automatic Tool AgentsZhengliang Shi, Shen Gao, Lingyong Yan, Yue Feng 等WWW 2025 · 被引用 59 次
- Group-Level Data Selection for Efficient PretrainingZichun Yu, Fei Peng, Jie Lei, Arnold Overwijk 等NeurIPS 2025 · 被引用 13 次
- InnovatorBench: Evaluating Agents' Ability to Conduct Innovative AI ResearchYunze Wu, Dayuan Fu, Weiye Si, Zhen Huang 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper25
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- SELECting over Tokens: Curating Pre-training Data at Scale via Token ClassificationXin Tong, Weidong Zhang, Jiaang Li, Haibin Chen 等ACL 2026
- LLaMA Pro: Progressive LLaMA with Block ExpansionChengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu 等ACL 2024
- DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language ModelsRanchi Zhao, Zhen Leng Thai, Yifan Zhang, Shengding Hu 等EMNLP 2024
- Rewriting Pre-Training Data Boosts LLM Performance in Math and CodeKazuki Fujii, Yukito Tajima, Sakae Mizuki, Masaki Kawamura 等ICLR 2026 · 被引用 21 次
- Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-trainingZheheng Luo, Xin Zhang, Xiao Liu, Haoling Li 等ACL 2025 · 被引用 8 次
