Are labels informative in semi-supervised learning? Estimating and leveraging the missing-data mechanism
Aude Sportisse, Hugo Schmutz, Olivier Humbert, Charles Bouveyron, Pierre-Alexandre Mattei
Abstract
Semi-supervised learning is a powerful technique for leveraging unlabeled data to improve machine learning models, but it can be affected by the presence of ``informative'' labels, which occur when some classes are more likely to be labeled than others. In the missing data literature, such labels are called missing not at random. In this paper, we propose a novel approach to address this issue by estimating the missing-data mechanism and using inverse propensity weighting to debias any SSL algorithm, including those using data augmentation. We also propose a likelihood ratio test to assess whether or not labels are indeed informative. Finally, we demonstrate the performance of the proposed methods on different datasets, in particular on two medical datasets for which we design pseudo-realistic missing data scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e646c35-3c41-472a-9ea1-6042a585a3c0Cited by top-tier papers2
- From Biased Selective Labels to Pseudo-Labels: An Expectation-Maximization Framework for Learning from Biased DecisionsTrenton Chang, Jenna WiensICML 2024 · 1 citation
- Efficient Causal Decision Making with One-sided FeedbackJianing Chu, Shu Yang, Wenbin Lu, Pulak GhoshICLR 2025
Builds on9
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 630 citations
- Open-World Semi-Supervised LearningKaidi Cao, Maria Brbic, Jure LeskovecICLR 2022 · 246 citations
- Safe Deep Semi-Supervised Learning for Unseen-Class Unlabeled DataLan-Zhe Guo, Zhenyu Zhang, Yuan Jiang, Yufeng Li et al.ICML 2020 · 243 citations
- Semi-Supervised Learning under Class Distribution MismatchYanbei Chen, Xiatian Zhu, Wei Li, Shaogang GongAAAI 2020 · 176 citations
Related papers
- On Non-Random Missing Labels in Semi-Supervised LearningXinting Hu, Yulei Niu, Chunyan Miao, Xian-Sheng Hua et al.ICLR 2022 · 22 citations
- Towards Semi-supervised Learning with Non-random Missing LabelsYue Duan, Zhen Zhao, Lei Qi, Luping Zhou et al.ICCV 2023 · 22 citations
- Data Augmentation with Diffusion for Open-Set Semi-Supervised LearningSeonghyun Ban, Heesan Kong, Kee-Eung KimNeurIPS 2024 · 4 citations
- Distribution Aligning Refinery of Pseudo-label for Imbalanced Semi-supervised LearningJaehyung Kim, Youngbum Hur, Sejun Park, Eunho Yang et al.NeurIPS 2020 · 209 citations
- ACPL: Anti-curriculum Pseudo-labelling for Semi-supervised Medical Image ClassificationFengbei Liu, Yu Tian, Yuanhong Chen, Yuyuan Liu et al.CVPR 2022 · 119 citations
