Leveraging Local Variance for Pseudo-Label Selection in Semi-supervised Learning
Zeping Min, Jinfeng Bai, Chengfei Li
Abstract
Semi-supervised learning algorithms that use pseudolabeling have become increasingly popular for improving model performance by utilizing both labeled and unlabeled data. In this paper, we offer a fresh perspective on the selection of pseudo-labels, inspired by theoretical insights. We suggest that pseudo-labels with a high degree of local variance are more prone to inaccuracies. Based on this premise, we introduce the Local Variance Match (LVM) method, which aims to optimize the selection of pseudo-labels in semi-supervised learning (SSL) tasks. Our methodology is validated through a series of experiments on widely-used image classification datasets, such as CIFAR-10, CIFAR-100, and SVHN, spanning various labeled data quantity scenarios. The empirical findings show that the LVM method substantially outpaces current SSL techniques, achieving stateof-the-art results in many of these scenarios. For instance, we observed an error rate of 5.41% on CIFAR-10 with a single label for each class, 35.87% on CIFAR-100 when using four labels per class, and 1.94% on SVHN with four labels for each class. Notably, the standout error rate of 5.41% is less than 1% shy of the performance in a fully-supervised learning environment. In experiments on ImageNet with 100k labeled data, the LVM also reached state-of-the-art outcomes. Additionally, the efficacy of the LVM method is further validated by its stellar performance in speech recognition experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb7548af-262d-46bb-89a7-53e6433a5bffCited by top-tier papers2
- VerifyMatch: A Semi-Supervised Learning Paradigm for Natural Language Inference with Confidence-Aware MixUpSeoyeon Park, Cornelia CarageaEMNLP 2024 · 1 citation
- Bypassing the Transport Plan: Dynamic Reweighting for Out-of-Distribution Detection with Optimal TransportYang Xiao, Weiming Liu, Jun Dan, Tengyue Xu et al.CVPR 2026
Builds on9
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu et al.NeurIPS 2021 · 1,389 citations
- ReMixMatch: Semi-Supervised Learning with Distribution Matching and Augmentation AnchoringDavid Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin et al.ICLR 2020 · 469 citations
- Dash: Semi-Supervised Learning with Dynamic ThresholdingYi Xu, Lei Shang, Jinxing Ye, Qi Qian et al.ICML 2021 · 287 citations
- AdaMatch: A Unified Approach to Semi-Supervised Learning and Domain AdaptationDavid Berthelot, Rebecca Roelofs, Kihyuk Sohn, Nicholas Carlini et al.ICLR 2022 · 180 citations
Related papers
- SoftMatch: Addressing the Quantity-Quality Tradeoff in Semi-supervised LearningHao Chen, Ran Tao, Yue Fan, Yidong Wang et al.ICLR 2023
- Int*-Match: Balancing Intra-Class Compactness and Inter-Class Discrepancy for Semi-Supervised Speaker RecognitionXingmei Wang, Jinghan Liu, Jiaxiang Meng, Boquan Li et al.AAAI 2025 · 1 citation
- Boosting Semi-Supervised Learning by Exploiting All Unlabeled DataYuhao Chen, Xin Tan, Borui Zhao, Zhaowei Chen et al.CVPR 2023
- ScaleMatch: Multi-scale Consistency Enhancement for Semi-supervised Semantic SegmentationLiang Lv, Lefei ZhangAAAI 2025 · 5 citations
- MarginMatch: Improving Semi-Supervised Learning with Pseudo-MarginsTiberiu Sosea, Cornelia CarageaCVPR 2023
