Diffusion Curriculum: Synthetic-to-Real Data Curriculum via Image-Guided Diffusion
Yijun Liang, Shweta Bhardwaj, Tianyi Zhou
摘要
Low-quality or scarce data has posed significant challenges for training deep neural networks in practice. While classical data augmentation cannot produce very different new data, diffusion models open up a new door to build selfevolving AI by generating high-quality and diverse synthetic data through text-guided prompts. However, text-only guidance cannot control synthetic images' proximity to the original images, resulting in out-of-distribution data detrimental to model performance. To overcome the limitation, we study image guidance to achieve a spectrum of interpolations between synthetic and real images. With stronger image guidance, the generated images are similar to the training data, but are hard to learn. With weaker image guidance, the synthetic images will be easier to learn but suffer from a larger distribution gap to the original data. The generated full spectrum of data enables us to build a novel “Diffusion CurricuLum (DisCL)”. DisCL adjusts the image guidance level of image synthesis for each training stage: It identifies and focuses on hard samples for the model and assesses the most effective guidance level of synthetic images to improve hard data learning. We apply DisCL to two challenging tasks: long-tail (LT) classification and learning from lowquality data. It focuses on lower-guidance images of high quality to learn prototypical features as a warm-up for learning higher-guidance images that might be weak on diversity or quality. DisCL achieves a gain of 2.7 % and 2.1 % in OOD and ID macro-accuracy when applied to iWildCam dataset. On ImageNet-LT, DisCL improves the base model's tail-class accuracy from 4.4 % to 23.64 % and leads to a 4.02 % improvement in all-class accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Sample-Efficient Multi-Round Generative Data Augmentation for Long-Tail Instance SegmentationByunghyun Kim, Minyoung Bae, Jae-Gil LeeNeurIPS 2025 · 被引用 3 次
- Resolving the Identity Crisis in Text-to-Image GenerationShubhankar Borse, Farzad Farhadzadeh, Munawar Hayat, Fatih PorikliCVPR 2026 · 被引用 2 次
- Risk-Bounded Distribution Reconstruction: Stable Statistic Calibration for Long-Tailed RecognitionGuanliang Liu, Wenchao Chen, Long Tian, Xuefei Cao 等ICML 2026
它引用的顶会 Paper21
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song 等ICLR 2022 · 被引用 2,128 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- DiffuLT: Diffusion for Long-tail Recognition Without External KnowledgeJie Shao, Ke Zhu, Hanxiao Zhang, Jianxin WuNeurIPS 2024 · 被引用 17 次
- Generating Images of Rare Concepts Using Pre-trained Diffusion ModelsDvir Samuel, Rami Ben-Ari, Simon Raviv, Nir Darshan 等AAAI 2024 · 被引用 82 次
- GOOD: Training-Free Guided Diffusion Sampling for Out-of-Distribution DetectionXin Gao, Jiyao Liu, Guanghao Li, Yueming Lyu 等NeurIPS 2025 · 被引用 9 次
- Generative Data Mining with Longtail-Guided DiffusionDavid S. Hayden, Mao Ye, Timur Garipov, Gregory P. Meyer 等ICML 2025
- Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based PerspectiveRui Huang, Shitong Shao, Zikai Zhou, Pukun Zhao 等CVPR 2026 · 被引用 7 次
