Towards Discovering the Effectiveness of Moderately Confident Samples for Semi-Supervised Learning
Hui Tang, Kui Jia
Abstract
Semi-supervised learning (SSL) has been studied for a long time to solve vision tasks in data-efficient application scenarios. SSL aims to learn a good classification model using a few labeled data together with large-scale unlabeled data. Recent advances achieve the goal by combining multiple SSL techniques, e.g., self-training and consistency regularization. From unlabeled samples, they usually adopt a confidence filter (CF) to select reliable ones with high prediction confidence. In this work, we study whether the moderately confident samples are useless and how to select the useful ones to improve model optimization. To answer these problems, we propose a novel Taylor expansion inspired filtration (TEIF) framework, which admits the samples of moderate confidence with similar feature or gradient to the respective one averaged over the labeled and highly confident unlabeled data. It can produce a stable and new information induced network update, leading to better generalization. Two novel filters are derived from this framework and can be naturally explained in two perspectives. One is gradient synchronization filter (GSF), which strengthens the optimization dynamic of fully-supervised learning; it selects the samples whose gradients are similar to class-wise majority gradients. The other is prototype proximity filter (PPF), which involves more prototypical samples in training to learn better semantic representations; it selects the samples near class-wise prototypes. They can be integrated into SSL methods with CF. We use the state-of-the-art Fix-Match as the baseline. Experiments on popular SSL benchmarks show that we achieve the new state of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5cfe467-0d77-4d52-a0c5-5a2d518ca1d8Cited by top-tier papers5
- Diffusion Models and Semi-Supervised Learners Benefit Mutually with Few LabelsZebin You, Yong Zhong, Fan Bao, Jiacheng Sun et al.NeurIPS 2023 · 61 citations
- Shrinking Class Space for Enhanced Certainty in Semi-Supervised LearningLihe Yang, Zhen Zhao, Lei Qi, Yu Qiao et al.ICCV 2023 · 27 citations
- Multi-Task Label Discovery via Hierarchical Task Tokens for Partially Annotated Dense PredictionsJingdong Zhang, Hanrong Ye, Xin Li, Wenping Wang et al.ACM MM 2025 · 1 citation
- RefTeacher: A Strong Baseline for Semi-Supervised Referring Expression ComprehensionJiamu Sun, Gen Luo, Yiyi Zhou, Xiaoshuai Sun et al.CVPR 2023
- Towards Calibrated Deep Clustering NetworkYuheng Jia, Jianhong Cheng, Hui Liu, Junhui HouICLR 2025
Builds on14
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu et al.NeurIPS 2021 · 1,389 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
Related papers
- RegMixMatch: Optimizing Mixup Utilization in Semi-Supervised LearningHaorong Han, Jidong Yuan, Chixuan Wei, Zhongyang YuAAAI 2025 · 7 citations
- Enhancing Sample Utilization through Sample Adaptive Augmentation in Semi-Supervised LearningGuan Gui, Zhen Zhao, Lei Qi, Luping Zhou et al.ICCV 2023 · 10 citations
- Time-Consistent Self-Supervision for Semi-Supervised LearningTianyi Zhou, Shengjie Wang, Jeff A. BilmesICML 2020 · 58 citations
- Dash: Semi-Supervised Learning with Dynamic ThresholdingYi Xu, Lei Shang, Jinxing Ye, Qi Qian et al.ICML 2021 · 287 citations
- HyperMatch: Noise-Tolerant Semi-Supervised Learning via Relaxed Contrastive ConstraintBeitong Zhou, Jing Lu, Kerui Liu, Yunlu Xu et al.CVPR 2023
