Data Augmentation with Diffusion for Open-Set Semi-Supervised Learning
Seonghyun Ban, Heesan Kong, Kee-Eung Kim
Abstract
Semi-supervised learning (SSL) seeks to utilize unlabeled data to overcome the limited amount of labeled data and improve model performance. However, many SSL methods typically struggle in real-world scenarios, particularly when there is a large number of irrelevant instances in the unlabeled data that do not belong to any class in the labeled data. Previous approaches often downweight instances from irrelevant classes to mitigate the negative impact of class distribution mismatch on model training. However, by discarding irrelevant instances, they may result in the loss of valuable information such as invariance, regularity, and diversity within the data. In this paper, we propose a data-centric generative augmentation approach that leverages a diffusion model to enrich labeled data using both labeled and unlabeled samples. A key challenge is extracting the diversity inherent in the unlabeled data while mitigating the generation of samples irrelevant to the labeled data. To tackle this issue, we combine diffusion model training with a discriminator that identifies and reduces the impact of irrelevant instances. We also demonstrate that such a trained diffusion model can even convert an irrelevant instance into a relevant one, yielding highly effective synthetic data for training. Through a comprehensive suite of experiments, we show that our data augmentation approach significantly enhances the performance of SSL methods, especially in the presence of class distribution mismatch.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf088b2f-b343-4725-982d-2074ede37217Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Open-World Semi-Supervised LearningKaidi Cao, Maria Brbic, Jure LeskovecICLR 2022 · 246 citations
- Safe Deep Semi-Supervised Learning for Unseen-Class Unlabeled DataLan-Zhe Guo, Zhenyu Zhang, Yuan Jiang, Yufeng Li et al.ICML 2020 · 243 citations
Related papers
- Towards Unbiased Learning in Semi-Supervised Semantic SegmentationRui Sun, Huayu Mai, Wangkai Li, Tianzhu ZhangICLR 2025
- MatchMask: Mask-Centric Generative Data Augmentation for Label-Scarce Semantic SegmentationYuqi Lin, Hao Zhang, Wenqi Shao, Shiqu Liu et al.CVPR 2026
- Exploring One-Shot Semi-supervised Federated Learning with Pre-trained Diffusion ModelsMingzhao Yang, Shangchao Su, Bin Li, Xiangyang XueAAAI 2024 · 52 citations
- Speech Self-Supervised Learning Using Diffusion Model Synthetic DataHeting Gao, Kaizhi Qian, Junrui Ni, Chuang Gan et al.ICML 2024 · 8 citations
- Towards Generic Semi-Supervised Framework for Volumetric Medical Image SegmentationHaonan Wang, Xiaomeng LiNeurIPS 2023 · 75 citations
