Diffusion Models and Semi-Supervised Learners Benefit Mutually with Few Labels
Zebin You, Yong Zhong, Fan Bao, Jiacheng Sun, Chongxuan Li, Jun Zhu
Abstract
In an effort to further advance semi-supervised generative and classification tasks, we propose a simple yet effective training strategy called dual pseudo training (DPT), built upon strong semi-supervised learners and diffusion models. DPT operates in three stages: training a classifier on partially labeled data to predict pseudo-labels; training a conditional generative model using these pseudo-labels to generate pseudo images; and retraining the classifier with a mix of real and pseudo images. Empirically, DPT consistently achieves SOTA performance of semi-supervised generation and classification across various settings. In particular, with one or two labels per class, DPT achieves a Fréchet Inception Distance (FID) score of 3.08 or 2.52 on ImageNet 256 × 256. Besides, DPT outperforms competitive semi-supervised baselines substantially on ImageNet classification tasks, achieving top-1 accuracies of 59.0 (+2.8), 69.5 (+3.0), and 74.4 (+2.0) with one, two, or five labels per class, respectively. Notably, our results demonstrate that diffusion can generate realistic images with only a few labels (e.g., < 0.1%) and generative augmentation remains viable for semi-supervised classification. Our code is available at https://github.com/ML-GSAI/DPT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5761010-e028-48a9-83a2-17a00471484bCited by top-tier papers21
- Dataset Diffusion: Diffusion-based Synthetic Data Generation for Pixel-Level Semantic SegmentationQuang Nguyen, Truong Vu, Anh Tran, Khoi NguyenNeurIPS 2023 · 154 citations
- Toward Understanding Generative Data AugmentationChenyu Zheng, Guoqiang Wu, Chongxuan LiNeurIPS 2023 · 51 citations
- ScaleLong: Towards More Stable Training of Diffusion Model via Scaling Network Long Skip ConnectionZhongzhan Huang, Pan Zhou, Shuicheng Yan, Liang LinNeurIPS 2023 · 41 citations
- Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion ModelMin Zhao, Hongzhou Zhu, Chendong Xiang, Kaiwen Zheng et al.NeurIPS 2024 · 33 citations
- MotionMix: Weakly-Supervised Diffusion for Controllable Motion GenerationNhat M. Hoang, Kehong Gong, Chuan Guo, Michael Bi MiAAAI 2024 · 11 citations
Builds on58
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based PerspectiveRui Huang, Shitong Shao, Zikai Zhou, Pukun Zhao et al.CVPR 2026 · 7 citations
- Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet ClonesMert Bülent Sariyildiz, Karteek Alahari, Diane Larlus, Yannis KalantidisCVPR 2023
- Meta Pseudo LabelsHieu Pham, Zihang Dai, Qizhe Xie, Quoc V. LeCVPR 2021
- Refining Generative Process with Discriminator Guidance in Score-based Diffusion ModelsDongjun Kim, Yeongmin Kim, Se Jung Kwon, Wanmo Kang et al.ICML 2023 · 109 citations
- PseudoSeg: Designing Pseudo Labels for Semantic SegmentationYuliang Zou, Zizhao Zhang, Han Zhang, Chun-Liang Li et al.ICLR 2021 · 364 citations
