Training Subset Selection for Weak Supervision
Hunter Lang, Aravindan Vijayaraghavan, David A. Sontag
Abstract
Existing weak supervision approaches use all the data covered by weak signals to train a classifier. We show both theoretically and empirically that this is not always optimal. Intuitively, there is a tradeoff between the amount of weakly-labeled data and the precision of the weak labels. We explore this tradeoff by combining pretrained data representations with the cut statistic (Muhlenbach et al., 2004) to select (hopefully) high-quality subsets of the weakly-labeled training data. Subset selection applies to any label model and classifier and is very simple to plug in to existing weak supervision pipelines, requiring just a few lines of code. We show our subset selection method improves the performance of weak supervision for a wide range of label models, classifiers, and datasets. Using less weakly-labeled data improves the accuracy of weak supervision pipelines by up to 19% (absolute) on benchmark tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Theoretical Analysis of Weak-to-Strong GeneralizationHunter Lang, David A. Sontag, Aravindan VijayaraghavanNeurIPS 2024 · 59 citations
- Neighborhood-Regularized Self-Training for Learning with Few LabelsRan Xu, Yue Yu, Hejie Cui, Xuan Kan et al.AAAI 2023 · 29 citations
- SLaM: Student-Label Mixing for Distillation with Unlabeled ExamplesVasilis Kontonis, Fotis Iliopoulos, Khoa Trinh, Cenk Baykal et al.NeurIPS 2023 · 10 citations
- Characterizing the Impacts of Semi-supervised Learning for Weak SupervisionJeffrey Li, Jieyu Zhang, Ludwig Schmidt, Alexander J. RatnerNeurIPS 2023 · 9 citations
- WISER: Weak Supervision and Supervised Representation Learning to Improve Drug Response Prediction in CancerKumar Shubham, Aishwarya Jayagopal, Syed Mohammed Danish, Prathosh A. P. et al.ICML 2024 · 8 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 182 citations
Related papers
- FastClass: A Time-Efficient Approach to Weakly-Supervised Text ClassificationTingyu Xia, Yue Wang, Yuan Tian, Yi ChangEMNLP 2022 · 1 citation
- Weaker Than You Think: A Critical Look at Weakly Supervised LearningDawei Zhu, Xiaoyu Shen, Marius Mosbach, Andreas Stephan et al.ACL 2023 · 13 citations
- Training Data Subset Selection for Regression with Controlled Generalization ErrorDurga Sivasubramanian, Rishabh K. Iyer, Ganesh Ramakrishnan, Abir DeICML 2021 · 25 citations
- The Perils of Learning From Unlabeled Data: Backdoor Attacks on Semi-supervised LearningVirat Shejwalkar, Lingjuan Lyu, Amir HoumansadrICCV 2023 · 15 citations
- Boosting Few-Shot Visual Learning With Self-SupervisionSpyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez et al.ICCV 2019 · 445 citations
