Improving the Scaling Laws of Synthetic Data with Deliberate Practice
Reyhane Askari Hemmat, Mohammad Pezeshki, Elvis Dohmatob, Florian Bordes, Pietro Astolfi, Melissa Hall, Jakob Verbeek, Michal Drozdzal, Adriana Romero-Soriano
摘要
Inspired by the principle of deliberate practice in human learning, we propose Deliberate Practice for Synthetic Data Generation (DP), a novel framework that improves sample efficiency through dynamic synthetic data generation. Prior work has shown that scaling synthetic data is inherently challenging, as naively adding new data leads to diminishing returns. To address this, pruning has been identified as a key mechanism for improving scaling, enabling models to focus on the most informative synthetic samples. Rather than generating a large dataset and pruning it afterward, DP efficiently approximates the direct generation of informative samples. We theoretically show how training on challenging, informative examples improves scaling laws and empirically validate that DP achieves better scaling performance with significantly fewer training samples and iterations. On ImageNet-100, DP generates 3.4× fewer samples and requires six times fewer iterations, while on ImageNet-1k, it generates 8× fewer samples with a 30% reduction in iterations, all while achieving superior performance compared to prior work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Inference-time Physics Alignment of Video Generative Models with Latent World ModelsJianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez 等CVPR 2026 · 被引用 32 次
- A Closer Look at Model Collapse: From a Generalization-to-Memorization PerspectiveLianghe Shi, Meng Wu, Huijie Zhang, Zekai Zhang 等NeurIPS 2025 · 被引用 22 次
- Teaching Models to Teach Themselves: Reasoning at the Edge of LearnabilityShobhita Sundaram, John Quan, Ariel Kwiatkowski, Kartik Ahuja 等ICML 2026 · 被引用 16 次
- Why Less is More (Sometimes): A Theory of Data CurationElvis Dohmatob, Mohammad Pezeshki, Reyhane Askari HemmatICLR 2026 · 被引用 11 次
- Increasing the Utility of Synthetic Images through Chamfer GuidanceNicola Dall'Asen, Xiaofeng Zhang, Reyhane Askari Hemmat, Melissa Hall 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper18
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Diversify Your Vision Datasets with Automatic Diffusion-based AugmentationLisa Dunlap, Alyssa Umino, Han Zhang, Jiezhi Yang 等NeurIPS 2023 · 被引用 136 次
相关 Paper
- OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning FrameworkChenhan Jin, Shengze Xu, Qingsong Wang, Fan JIA 等ICLR 2026 · 被引用 3 次
- Real-Fake: Effective Training Data Synthesis Through Distribution MatchingJianhao Yuan, Jie Zhang, Shuyang Sun, Philip Torr 等ICLR 2024 · 被引用 47 次
- Scale Efficient Training for Large DatasetsQing Zhou, Junyu Gao, Qi WangCVPR 2025
- UNSEEN: Enhancing Dataset Pruning from a Generalization PerspectiveFurui Xu, Shaobo Wang, Jiajun Zhang, Chenghao Sun 等AAAI 2026
- Synthetic Data Can Also Teach: Synthesizing Effective Data for Unsupervised Visual Representation LearningYawen Wu, Zhepeng Wang, Dewen Zeng, Yiyu Shi 等AAAI 2023 · 被引用 20 次
