Re-Evaluating the Impact of Unseen-Class Unlabeled Data on Semi-Supervised Learning Model
Rundong He, Yicong Dong, Lanzhe Guo, Yilong Yin, Tailin Wu
Abstract
Semi-supervised learning (SSL) effectively leverages unlabeled data and has been proven successful across various fields. Current safe SSL methods believe that unseen classes in unlabeled data harm the performance of SSL models. However, previous methods for assessing the impact of unseen classes on SSL model performance are flawed. They fix the size of the unlabeled dataset and adjust the proportion of unseen classes within the unlabeled data to assess the impact. This process contravenes the principle of controlling variables. Adjusting the proportion of unseen classes in unlabeled data alters the proportion of seen classes, meaning the decreased classification performance of seen classes may not be due to an increase in unseen class samples in the unlabeled data, but rather a decrease in seen class samples. Thus, the prior flawed assessment standard that "unseen classes in unlabeled data can damage SSL model performance" may not always hold true. This paper strictly adheres to the principle of controlling variables, maintaining the proportion of seen classes in unlabeled data while only changing the unseen classes across five critical dimensions, to investigate their impact on SSL models from global robustness and local robustness. Experiments demonstrate that unseen classes in unlabeled data do not necessarily impair the performance of SSL models; in fact, under certain conditions, unseen classes may even enhance them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-LearningTrinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen et al.KDD 2026 · 1 citation
- Towards Out-of-Modal Generalization without Instance-level Modal CorrespondenceZhuo Huang, Gang Niu, Bo Han, Masashi Sugiyama et al.ICLR 2025
Builds on17
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- ReMixMatch: Semi-Supervised Learning with Distribution Matching and Augmentation AnchoringDavid Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin et al.ICLR 2020 · 469 citations
- CoMatch: Semi-supervised Learning with Contrastive Graph RegularizationJunnan Li, Caiming Xiong, Steven C. H. HoiICCV 2021 · 333 citations
- Open-World Semi-Supervised LearningKaidi Cao, Maria Brbic, Jure LeskovecICLR 2022 · 246 citations
Related papers
- Safe Deep Semi-Supervised Learning for Unseen-Class Unlabeled DataLan-Zhe Guo, Zhenyu Zhang, Yuan Jiang, Yufeng Li et al.ICML 2020 · 243 citations
- Not All Parameters Should Be Treated Equally: Deep Safe Semi-supervised Learning under Class Distribution MismatchRundong He, Zhongyi Han, Yang Yang, Yilong YinAAAI 2022 · 32 citations
- Robust Semi-Supervised Learning when Not All Classes have LabelsLan-Zhe Guo, Yi-Ge Zhang, Zhi-Fan Wu, Jie-Jing Shao et al.NeurIPS 2022 · 63 citations
- Safe-Student for Safe Deep Semi-Supervised Learning with Unseen-Class Unlabeled DataRundong He, Zhongyi Han, Xiankai Lu, Yilong YinCVPR 2022 · 49 citations
- Rethinking Safe Semi-supervised Learning: Transferring the Open-set Problem to A Close-set OneQiankun Ma, Jiyao Gao, Bo Zhan, Yunpeng Guo et al.ICCV 2023 · 15 citations
