Scalable Data Synthesis through Human-like Cognitive Imitation and Data Recombination
Zhongyi Ye, Weitai Zhang, Xinyuan Zhou, Yongxin Zhu, Ninghui Rao, Enhong Chen
摘要
Large language models (LLMs) rely on massive amounts of training data, however, the quantity of empirically observed data is limited. To alleviate this issue, lots of LLMs leverage synthetic data to enhance the quantity of training data. Despite significant advancements in LLMs, the efficiency and scalability characteristics of data synthesis during pre-training phases remain insufficiently explored. In this work, we propose a novel data synthesis framework, Cognitive Combination Synthesis (CCS), designed to achieve highly efficient and scalable data synthesis. Specifically, our methodology mimics human cognitive behaviors by re-combining and interconnecting heterogeneous data from diverse sources thereby enhancing advanced reasoning capabilities in LLMs. Extensive experiments demonstrate that: (1) effective data organization is essential, and our mapping-based combination learning approach significantly improves data utilization efficiency; (2) by enhancing data diversity , accuracy , and complexity , our synthetic data scales beyond 100B tokens, revealing CCS’s strong scalability. Our findings highlight the impact of data organization methods on LLM learning efficiency and the significant potential of scalable synthetic data to enhance model reasoning capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu 等ICLR 2024 · 被引用 637 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- MAmmoTH2: Scaling Instructions from the WebXiang Yue, Tianyu Zheng, Ge Zhang, Wenhu ChenNeurIPS 2024 · 被引用 176 次
- MathScale: Scaling Instruction Tuning for Mathematical ReasoningZhengyang Tang, Xingxing Zhang, Benyou Wang, Furu WeiICML 2024 · 被引用 163 次
相关 Paper
- SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and BeyondJunteng Liu, Yuanxiang Fan, Zhuo Jiang, Han Ding 等NeurIPS 2025 · 被引用 49 次
- How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data CompositionGuanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li 等ACL 2024 · 被引用 39 次
- Composition-Grounded Data Synthesis for Visual ReasoningXinyi Gu, Jiayuan Mao, Zhang-Wei Hong, Zhuoran Yu 等ICLR 2026 · 被引用 1 次
- Learning from Synthetic Data Improves Multi-hop ReasoningAnmol Kabra, Yilun Yin, Albert Gong, Kamilė Stankevičiūtė 等ICLR 2026 · 被引用 6 次
- MIND: Math Informed syNthetic Dialogues for Pretraining LLMsSyeda Nahida Akter, Shrimai Prabhumoye, John Kamalu, Sanjeev Satheesh 等ICLR 2025
