Rewriting Pre-Training Data Boosts LLM Performance in Math and Code
Kazuki Fujii, Yukito Tajima, Sakae Mizuki, Masaki Kawamura, Hinari Shimada, Taihei Shiotani, Koshiro Saito, Masanari Oi, Taishi Nakamura, Takumi Okamoto, Shigeki Ishida, Kakeru Hattori
Abstract
The performance of large language models (LLMs) in program synthesis and mathematical reasoning is fundamentally limited by the quality of their pre-training corpora.
We introduce two openly licensed pre-training datasets, released under the Llama 3.3 Community License, that significantly enhance LLM performance by systematically rewriting public data. SwallowCode (16.1 billion tokens) refines Python snippets from The-Stack-v2 through a novel four-stage pipeline: syntax validation, pylint-based style filtering, and a two-stage LLM rewriting process that enforces style conformity and transforms snippets into self-contained, algorithmically efficient examples. Unlike prior methods that rely on exclusionary filtering or limited transformations, our transform-and-retain approach refines low-quality code, maximizing data utility.
SwallowMath (2.3 billion tokens) enhances Finemath-4+ by removing boilerplate, restoring context, and reformatting solutions into concise, step-by-step explanations. Within a fixed 50 billion token training budget, continual pre-training of Llama-3.1-8B with SwallowCode boosts pass@1 by +17.0 on HumanEval and +16.1 on HumanEval+ compared to Stack-Edu, surpassing the baseline model's code generation capabilities. Similarly, substituting SwallowMath yields +12.4 accuracy on GSM8K and +7.6 on MATH. Ablation studies confirm that each pipeline stage contributes incrementally, with rewriting yielding the largest gains.
By releasing datasets, prompts, checkpoints, and pipeline code, we ensure reproducibility and provide a transferable transform-and-retain methodology that can be adapted to other base models and LLM rewriting setups.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3d727cf-49f5-49e2-b391-5c0750c91dcbCited by top-tier papers5
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture SearchYuxian Gu, Qinghao Hu, Haocheng Xi, Junyu Chen et al.NeurIPS 2025 · 39 citations
- Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in SpeedYonggan Fu, Lexington Whalen, Zhifan Ye, Xin Dong et al.ICML 2026 · 22 citations
- Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-trainingShengrui Li, Fei zhao, Kaiyan Zhao, Jieying Ye et al.ICML 2026 · 3 citations
- Proxy Compression for Language ModelingLin Zheng, Li Xinyu, Qian Liu, Xiachong Feng et al.ICML 2026 · 3 citations
- Droid: A Resource Suite for AI-Generated Code DetectionDaniil Orel, Indraneil Paul, Iryna Gurevych, Preslav NakovEMNLP 2025
Builds on7
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web TextKeiran Paster, Marco Dos Santos, Zhangir Azerbayev, Jimmy BaICLR 2024 · 140 citations
- LLM-Assisted Code Cleaning For Training Accurate Code GeneratorsNaman Jain, Tianjun Zhang, Wei-Lin Chiang, Joseph E. Gonzalez et al.ICLR 2024 · 49 citations
- Rephrasing the Web: A Recipe for Compute and Data-Efficient Language ModelingPratyush Maini, Skyler Seto, Richard He Bai, David Grangier et al.ACL 2024 · 13 citations
Related papers
- MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical CodeZimu Lu, Aojun Zhou, Ke Wang, Houxing Ren et al.ICLR 2025
- OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction DataShubham Toshniwal, Wei Du, Ivan Moshkov, Branislav Kisacanin et al.ICLR 2025
- MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical ReasoningShuo Yin, Weihao You, Zhilong Ji, Guoqiang Zhong et al.EMNLP 2024 · 3 citations
- Chain of Execution Supervision Promotes General Reasoning in Large Language ModelsNuo Chen, Zehua Li, Keqin Bao, Junyang Lin et al.NeurIPS 2025 · 6 citations
- Programming Every Example: Lifting Pre-training Data Quality Like Experts at ScaleFan Zhou, Zengzhi Wang, Qian Liu, Junlong Li et al.ICML 2025
