Don't fear the unlabelled: safe semi-supervised learning via debiasing
Hugo Schmutz, Olivier Humbert, Pierre-Alexandre Mattei
Abstract
Semi-supervised learning (SSL) provides an effective means of leveraging unlabelled data to improve a model's performance. Even though the domain has received a considerable amount of attention in the past years, most methods present the common drawback of lacking theoretical guarantees. Our starting point is to notice that the estimate of the risk that most discriminative SSL methods minimise is biased, even asymptotically. This bias impedes the use of standard statistical learning theory and can hurt empirical performance. We propose a simple way of removing the bias. Our debiasing approach is straightforward to implement and applicable to most deep SSL methods. We provide simple theoretical guarantees on the trustworthiness of these modified methods, without having to rely on the strong assumptions on the data distribution that SSL theory usually requires. In particular, we provide generalisation error bounds for the proposed methods. We evaluate debiased versions of different existing SSL methods, such as the Pseudolabel method and Fixmatch, and show that debiasing can compete with classic deep SSL techniques in various settings by providing better calibrated models. Additionally, we provide a theoretical explanation of the intuition of the popular SSL methods. An implementation of a debiased version of Fixmatch is available at https://github.com/HugoSchmutz/DeFixmatch
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c644b206-6955-4ee9-9a27-bd4f08ce0b24Cited by top-tier papers5
- SemiVisBooster: Boosting Semi-Supervised Learning for Fine-Grained Classification through Pseudo-Label Semantic GuidanceWenjin Zhang, Xinyu Li, Chenyang Gao, Ivan MarsicICCV 2025 · 4 citations
- Revisiting Active Sequential Prediction-Powered Mean EstimationMaria-Eleni Sfyraki, Jun-Kun WangICLR 2026 · 4 citations
- LCGC: Learning from Consistency Gradient Conflicting for Class-Imbalanced Semi-Supervised DebiasingWeiwei Xing, Yue Cheng, Hongzhu Yi, Xiaohui Gao et al.AAAI 2025 · 3 citations
- Enhancing Semi-Supervised Learning via Representative and Diverse Sample SelectionQian Shao, Jiangrui Kang, Qiyuan Chen, Zepeng Li et al.NeurIPS 2024 · 3 citations
- Towards Understanding Why FixMatch Generalizes Better Than Supervised LearningJingyang Li, Jiachun Pan, Vincent Y. F. Tan, Kim-Chuan Toh et al.ICLR 2025
Builds on13
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu et al.NeurIPS 2021 · 1,389 citations
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 630 citations
Related papers
- Debiased Self-Training for Semi-Supervised LearningBaixu Chen, Junguang Jiang, Ximei Wang, Pengfei Wan et al.NeurIPS 2022 · 162 citations
- CaliMatch: Adaptive Calibration for Improving Safe Semi-Supervised LearningJinsoo Bae, Seoung Bum Kim, Hyungrok DoICCV 2025 · 1 citation
- SoftMatch: Addressing the Quantity-Quality Tradeoff in Semi-supervised LearningHao Chen, Ran Tao, Yue Fan, Yidong Wang et al.ICLR 2023
- Boosting Semi-Supervised Learning by Exploiting All Unlabeled DataYuhao Chen, Xin Tan, Borui Zhao, Zhaowei Chen et al.CVPR 2023
- FlatMatch: Bridging Labeled Data and Unlabeled Data with Cross-Sharpness for Semi-Supervised LearningZhuo Huang, Li Shen, Jun Yu, Bo Han et al.NeurIPS 2023 · 50 citations
