Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling
Pratyush Maini, Skyler Seto, Richard He Bai, David Grangier, Yizhe Zhang, Navdeep Jaitly
Abstract
Large language models are trained on massive scrapes of the web, which are often unstructured, noisy, and poorly phrased. Current scaling laws show that learning from such data requires an abundance of both compute and data, which grows with the size of the model being trained. This is infeasible both because of the large compute costs and duration associated with pre-training, and the impending scarcity of high-quality data on the web. In this work, we propose Web Rephrase Augmented Pre-training (WRAP) that uses an off-the-shelf instruction-tuned model prompted to paraphrase documents on the web in specific styles such as "like Wikipedia" or in "questionanswer format" to jointly pre-train LLMs on real and synthetic rephrases. First, we show that using WRAP on the C4 dataset, which is naturally noisy, speeds up pre-training by ∼ 3×. At the same pre-training compute budget, it improves perplexity by more than 50% on average across different subsets of the Pile, and improves zero-shot question answer accuracy across 13 tasks by more than 2%. Second, we investigate the impact of the re-phrasing style on the performance of the model, offering insights into how the composition of the training data can impact the performance of LLMs in OOD settings. Our gains are attributed to the fact that re-phrased synthetic data (i) incorporates style diversity that closely reflects downstream evaluation style, and (ii) has higher 'quality' than web-scraped data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers55
- C3oT: Generating Shorter Chain-of-Thought Without Compromising EffectivenessYu Kang, Xianghui Sun, Liangyu Chen, Wei ZouAAAI 2025 · 162 citations
- MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence ModelsZichun Yu, Spandan Das, Chenyan XiongNeurIPS 2024 · 117 citations
- How to train data-efficient LLMsNoveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang, Jianmo Ni et al.ICLR 2026 · 106 citations
- Self-Adapting Language ModelsAdam Zweiger, Jyothish Pari, Han Guo, Yoon Kim et al.NeurIPS 2025 · 78 citations
- Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM ReasoningJaehun Jung, Seungju Han, Ximing Lu, Skyler Hallinan et al.NeurIPS 2025 · 50 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
Related papers
- RePro: Training Language Models to Faithfully Recycle the Web for PretrainingZichun Yu, Chenyan XiongICML 2026 · 2 citations
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and PitfallsFeiyang Kang, Newsha Ardalani, Michael Kuchnik, Youssef Emad et al.EMNLP 2025
- Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training EfficiencyEhsan Doostmohammadi, Marco KuhlmannEMNLP 2025
- Reusing Pre-Training Data at Test Time is a Compute MultiplierAlex Fang, Thomas Voice, Ruoming Pang, Ludwig Schmidt et al.ICLR 2026 · 4 citations
- Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language ModelsYukun Huang, Sanxing Chen, Jian Pei, Manzil Zaheer et al.ICLR 2026 · 1 citation
