Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
Rabeeh Karimi Mahabadi, Sanjeev Satheesh, Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro
Abstract
Pretraining large language models (LLMs) on high-quality, structured data such as mathematics and code substantially enhances reasoning capabilities. However, existing math-focused datasets built from Common Crawl suffer from degraded quality due to brittle extraction heuristics, lossy HTML-to-text conversion, and the failure to reliably preserve mathematical structure. In this work, we introduce Nemotron-CC-Math, a large-scale, high-quality mathematical corpus constructed from Common Crawl using a novel, domain-agnostic pipeline specifically designed for robust scientific text extraction. Unlike previous efforts, our pipeline recovers math across various formats (e.g., MathJax, KaTeX, MathML) by leveraging layout-aware rendering with lynx and a targeted LLM-based cleaning stage. This approach preserves the structural integrity of equations and code blocks while removing boilerplate, standardizing notation into L A T E X representation, and correcting inconsistencies. We collected a large, high-quality math corpus, namely Nemotron-CC-Math-3+ (133B tokens) and Nemotron-CC-Math-4+ (52B tokens). Notably, Nemotron-CC-Math-4+ not only surpasses all prior open math datasets-including Mega-Math, FineMath, and OpenWebMath-but also contains 5.5× more tokens than FineMath-4+, which was previously the highest-quality math pretraining dataset. When used to pretrain a Nemotron-T 8B model, our corpus yields +4.8 to +12.6 gains on MATH and +4.6 to +14.3 gains on MBPP+ over strong baselines, while also improving general-domain performance on MMLU and MMLU-Stem. We present the first pipeline to reliably extract scientific content-including math-from noisy web-scale data, yielding measurable gains in math, code, and general reasoning, and setting a new state of the art among open math pretraining corpora. To support open-source efforts, we release our code 1 and datasets 2 . * Rabeeh and Sanjeev are the primary authors and contributed equally.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd86dd52-d012-402f-a1ee-4cde00860857Cited by top-tier papers5
- STEM: Scaling Transformers with Embedding ModulesRanajoy Sadhukhan, Sheng Cao, Harry Dong, Changsheng Zhao et al.ICLR 2026 · 14 citations
- Understanding Dynamic Compute Allocation in Recurrent TransformersIbraheem Muhammad Moosa, Suhas Lohit, Ye Wang, Moitreya Chatterjee et al.ICML 2026 · 5 citations
- SEDD: Scalable and Efficient Dataset Deduplication with GPUsYoungjun Son, Chaewon Kim, Jaejin LeeKDD 2026 · 3 citations
- From Growing to Looping: A Unified View of Iterative Computation in LLMsFerdinand Kapl, Emmanouil Angelis, Kaitlin Maile, Johannes von Oswald et al.ICML 2026 · 2 citations
- WIST: Web-Grounded Iterative Self-Play Tree for Domain-Targeted Reasoning ImprovementFangyuan Li, Pengfei Li, Shijie Wang, Junqi Gao et al.ACL 2026
Builds on8
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
Related papers
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web TextKeiran Paster, Marco Dos Santos, Zhangir Azerbayev, Jimmy BaICLR 2024 · 140 citations
- MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical CodeZimu Lu, Aojun Zhou, Ke Wang, Houxing Ren et al.ICLR 2025
- Rewriting Pre-Training Data Boosts LLM Performance in Math and CodeKazuki Fujii, Yukito Tajima, Sakae Mizuki, Masaki Kawamura et al.ICLR 2026 · 21 citations
- Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining DatasetDan Su, Kezhi Kong, Ying Lin, Joseph Jennings et al.ACL 2025
- MathScale: Scaling Instruction Tuning for Mathematical ReasoningZhengyang Tang, Xingxing Zhang, Benyou Wang, Furu WeiICML 2024 · 163 citations
