SELECting over Tokens: Curating Pre-training Data at Scale via Token Classification
Xin Tong, Weidong Zhang, Jiaang Li, Haibin Chen, Shilei Liu, Langming Liu, Kangtao Lv, Yujin Yuan, Wenbo Su, Bo Zheng
摘要
The quality of pre-training data critically impacts the capabilities of large language models. Existing pipelines rely on expert-crafted heuristic rules, which primarily operate at the sample level and are based on coarse statistical indicators, thus lacking content-aware, fine-grained noise detection. While recent generative approaches, e.g., PROX-C, enable token-level refinement, their reliance on synthesizing Python code incurs prohibitive computational cost at scale and can introduce hallucinations into the refined data. To overcome these limitations, we propose SELECting over Tokens (SELECT), a novel framework that reframes data refinement as a highly efficient token classification task. SELECT classifies each token as either informative or noisy and subsequently removes the latter. This design achieves fine-grained data optimization while avoiding the inefficiency of generation, ensuring scalability. When evaluated on diverse downstream benchmarks, the model trained on SELECT-refined corpora, on average, outperforms the one trained on raw data by over 2% and exceeds the best heuristic baselines by more than 1% while preserving 17% more tokens than the latter. Furthermore, SELECT achieves higher average performance than the generative PROX-C across all experimental settings, and is 2.5x faster at inference, even with twice the parameters. Our results establish SELECT as an effective, efficient, and scalable solution for pre-training data optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller 等ACL 2025 · 被引用 552 次
- Rephrasing the Web: A Recipe for Compute and Data-Efficient Language ModelingPratyush Maini, Skyler Seto, Richard He Bai, David Grangier 等ACL 2024 · 被引用 13 次
- Programming Every Example: Lifting Pre-training Data Quality Like Experts at ScaleFan Zhou, Zengzhi Wang, Qian Liu, Junlong Li 等ICML 2025
相关 Paper
- Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-TuningJinlong Pang, Na Di, Zhaowei Zhu, Jiaheng Wei 等ICML 2025
- Explainable Token-level Noise Filtering for LLM Fine-tuning DatasetsYuchen Yang, Wenze Lin, Enhao Huang, Zhixuan Chu 等ICLR 2026 · 被引用 1 次
- Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For PerplexityYeongbin Seo, Gayoung Kim, Jaehyung Kim, Jinyoung YeoICLR 2026
- OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every IterationShaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu 等ICML 2026 · 被引用 12 次
- Predictive Data Selection: The Data That Predicts Is the Data That TeachesKaShun Shum, Yuzhen Huang, Hongjian Zou, Qi Ding 等ICML 2025
