Characterizing the Impacts of Semi-supervised Learning for Weak Supervision
Jeffrey Li, Jieyu Zhang, Ludwig Schmidt, Alexander J. Ratner
Abstract
Labeling training data is a critical and expensive step in producing high accuracy ML models, whether training from scratch or fine-tuning. To make labeling more efficient, two major approaches are programmatic weak supervision (WS) and semisupervised learning (SSL). More recent works have either explicitly or implicitly used techniques at their intersection, but in various complex and ad hoc ways. In this work, we define a simple, modular design space to study the use of SSL techniques for WS more systematically. Surprisingly, we find that fairly simple methods from our design space match the performance of more complex state-of-the-art methods, averaging a 3 p.p. increase in accuracy/F1-score across 8 standard WS benchmarks. Further, we provide practical guidance on when different components are worth their added complexity and training costs. Contrary to current understanding, we find SSL is not necessary to obtain the best performance on most existing WS benchmarks but is more effective when: (1) end models are smaller, and (2) WS provides labels for only a small portion of training examples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd0f008d-d724-48ea-b439-c9f29aacd694Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu et al.ICLR 2022 · 338 citations
- Learning from Rules Generalizing Labeled ExemplarsAbhijeet Awasthi, Sabyasachi Ghosh, Rasna Goyal, Sunita SarawagiICLR 2020 · 93 citations
- MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesJack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie et al.ACL 2023 · 88 citations
Related papers
- The Perils of Learning From Unlabeled Data: Backdoor Attacks on Semi-supervised LearningVirat Shejwalkar, Lingjuan Lyu, Amir HoumansadrICCV 2023 · 15 citations
- Learning Hyper Label Model for Programmatic Weak SupervisionRenzhi Wu, Shen-En Chen, Jieyu Zhang, Xu ChuICLR 2023 · 2 citations
- WeShap: Weak Supervision Source Evaluation with Shapley ValuesNaiqing Guan, Nick KoudasVLDB 2025
- Self-Tuning for Data-Efficient Deep LearningXimei Wang, Jinghan Gao, Mingsheng Long, Jianmin WangICML 2021 · 79 citations
- DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled SamplesYi Xu, Jiandong Ding, Lu Zhang, Shuigeng ZhouNeurIPS 2021 · 34 citations
