Can semi-supervised learning use all the data effectively? A lower bound perspective
Alexandru Tifrea, Gizem Yüce, Amartya Sanyal, Fanny Yang
Abstract
Prior works have shown that semi-supervised learning algorithms can leverage unlabeled data to improve over the labeled sample complexity of supervised learning (SL) algorithms. However, existing theoretical analyses focus on regimes where the unlabeled data is sufficient to learn a good decision boundary using unsupervised learning (UL) alone. This begs the question: Can SSL algorithms simultaneously improve upon both UL and SL? To this end, we derive a tight lower bound for 2-Gaussian mixture models that explicitly depends on the labeled and the unlabeled dataset size as well as the signal-to-noise ratio of the mixture distribution. Surprisingly, our result implies that no SSL algorithm can improve upon the minimax-optimal statistical error rates of SL or UL algorithms for these distributions. Nevertheless, we show empirically on real-world data that SSL algorithms can still outperform UL and SL methods. Therefore, our work suggests that, while proving performance gains for SSL algorithms is possible, it requires careful tracking of constants. * Equal contribution. Presented at the 37th Conference on Neural Information Processing Systems (NeurIPS 2023). 1 By error of UL we mean the prediction error up to sign. We formalize this paradigm of using UL first and then identifying the correct sign as UL+ in Section 2.2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b221f29f-9738-496d-a4b6-9fe13286616bCited by top-tier papers4
- Evaluating multiple models using labeled and unlabeled dataDivya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John V. Guttag et al.NeurIPS 2025 · 9 citations
- Semi-Supervised Sparse Gaussian Classification: Provable Benefits of Unlabeled DataEyar Azar, Boaz NadlerNeurIPS 2024 · 5 citations
- On the sample complexity of semi-supervised multi-objective learningTobias Wegel, Geelon So, Junhyung Park, Fanny YangNeurIPS 2025 · 3 citations
- Towards Understanding Why FixMatch Generalizes Better Than Supervised LearningJingyang Li, Jiachun Pan, Vincent Y. F. Tan, Kim-Chuan Toh et al.ICLR 2025
Builds on10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
Related papers
- Class-Imbalanced Semi-Supervised Learning with Adaptive ThresholdingLan-Zhe Guo, Yufeng LiICML 2022 · 148 citations
- The Perils of Learning From Unlabeled Data: Backdoor Attacks on Semi-supervised LearningVirat Shejwalkar, Lingjuan Lyu, Amir HoumansadrICCV 2023 · 15 citations
- FlatMatch: Bridging Labeled Data and Unlabeled Data with Cross-Sharpness for Semi-Supervised LearningZhuo Huang, Li Shen, Jun Yu, Bo Han et al.NeurIPS 2023 · 50 citations
- Towards Cost-Effective Learning: A Synergy of Semi-Supervised and Active LearningTianxiang Yin, Ningzhong Liu, Han SunCVPR 2025
- Towards Realistic Model Selection for Semi-supervised LearningMuyang Li, Xiaobo Xia, Runze Wu, Fengming Huang et al.ICML 2024 · 2 citations
