How Does Semi-supervised Learning with Pseudo-labelers Work? A Case Study
Yiwen Kou, Zixiang Chen, Yuan Cao, Quanquan Gu
摘要
Semi-supervised learning is a popular machine learning paradigm that utilizes a large amount of unlabeled data as well as a small amount of labeled data to facilitate learning tasks. While semi-supervised learning has achieved great success in training neural networks, its theoretical understanding remains largely open. In this paper, we aim to theoretically understand a semi-supervised learning approach based on pre-training and linear probing. We prove that, under a certain data generation model and two-layer convolutional neural network, the semi-supervised learning approach can achieve nearly zero test loss, while a neural network directly trained by supervised learning on the same amount of labeled data can only achieve constant test loss. Through this case study, we demonstrate a separation between semi-supervised learning and supervised learning in terms of test loss provided the same amount of labeled data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsZixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji 等ICML 2024 · 被引用 527 次
- Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context LearningDake Bu, Wei Huang, Andi Han, Atsushi Nitanda 等NeurIPS 2024 · 被引用 11 次
它引用的顶会 Paper5
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
- Benign Overfitting in Two-layer Convolutional Neural NetworksYuan Cao, Zixiang Chen, Misha Belkin, Quanquan GuNeurIPS 2022 · 被引用 121 次
- Feature Purification: How Adversarial Training Performs Robust Deep LearningZeyuan Allen-Zhu, Yuanzhi LiFOCS 2021 · 被引用 83 次
相关 Paper
- Self-Supervised Wasserstein Pseudo-Labeling for Semi-Supervised Image ClassificationFariborz Taherkhani, Ali Dabouei, Sobhan Soleymani, Jeremy M. Dawson 等CVPR 2021
- All Labels Are Not Created Equal: Enhancing Semi-Supervision via Label Grouping and Co-TrainingIslam Nassar, Samitha Herath, Ehsan Abbasnejad, Wray L. Buntine 等CVPR 2021
- SemPPL: Predicting Pseudo-Labels for Better Contrastive RepresentationsMatko Bosnjak, Pierre Harvey Richemond, Nenad Tomasev, Florian Strub 等ICLR 2023 · 被引用 4 次
- Curriculum Labeling: Revisiting Pseudo-Labeling for Semi-Supervised LearningPaola Cascante-Bonilla, Fuwen Tan, Yanjun Qi, Vicente OrdonezAAAI 2021 · 被引用 362 次
- Generalized Semi-Supervised Learning via Self-Supervised Feature AdaptationJiachen Liang, Ruibing Hou, Hong Chang, Bingpeng Ma 等NeurIPS 2023 · 被引用 7 次
