SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training
Nan He, Weichen Xiong, Hanwen Liu, Yi Liao, Lei Ding, Kai Zhang, Guohua Tang, Xiao Han, Yang Wei
摘要
The effectiveness of large language models (LLMs) is often hindered by duplicated data in their extensive pre-training datasets. Current approaches primarily focus on detecting and removing duplicates, which risks the loss of valuable information and neglects the varying degrees of duplication. To address this, we propose a soft deduplication method that maintains dataset integrity while selectively reducing the sampling weight of data with high commonness. Central to our approach is the concept of "data commonness", a metric we introduce to quantify the degree of duplication by measuring the occurrence probabilities of samples using an n-gram model. Empirical analysis shows that this method significantly improves training efficiency, achieving comparable perplexity scores with at least a 26% reduction in required training steps. Additionally, it enhances average few-shot downstream accuracy by 1.77% when trained for an equivalent duration. Importantly, this approach consistently improves performance, even on rigorously deduplicated datasets, indicating its potential to complement existing methods and become a standard pretraining process for LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting ObjectivesXiaoqing Zhang, Ang Lv, Yuhan Liu, Flood Sung 等ACL 2025 · 被引用 9 次
- Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models PretrainingPing Guo, Yubing Ren, Binbin Liu, Fengze Liu 等NeurIPS 2025 · 被引用 2 次
- FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language ModelsPukang Ye, Junwei Luo, Jiachen Shen, Saipan Zhou 等NeurIPS 2025 · 被引用 1 次
- Principled Synthetic Data Enables the First Scaling Laws for LLMs in RecommendationBenyu Zhang, Qiang Zhang, Jianpeng Cheng, Hong-You Chen 等ICML 2026
它引用的顶会 Paper2
相关 Paper
- Scalable Data Ablation Approximations for Language Models through Modular Training and MergingClara Na, Ian Magnusson, Ananya Harsh Jha, Tom Sherborne 等EMNLP 2024 · 被引用 2 次
- Datasets, Documents, and Repetitions: The Practicalities of Unequal Data QualityAlex Fang, Hadi Pouransari, Matt Jordan, Alexander Toshev 等NeurIPS 2025 · 被引用 6 次
- Dynamic Loss-Based Sample Reweighting for Improved Large Language Model PretrainingDaouda Sow, Herbert Woisetschläger, Saikiran Bulusu, Shiqiang Wang 等ICLR 2025
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 被引用 383 次
- Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language ModelsXinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang 等ACL 2025 · 被引用 15 次
