Training Subset Selection for Weak Supervision
Hunter Lang, Aravindan Vijayaraghavan, David A. Sontag
摘要
Existing weak supervision approaches use all the data covered by weak signals to train a classifier. We show both theoretically and empirically that this is not always optimal. Intuitively, there is a tradeoff between the amount of weakly-labeled data and the precision of the weak labels. We explore this tradeoff by combining pretrained data representations with the cut statistic (Muhlenbach et al., 2004) to select (hopefully) high-quality subsets of the weakly-labeled training data. Subset selection applies to any label model and classifier and is very simple to plug in to existing weak supervision pipelines, requiring just a few lines of code. We show our subset selection method improves the performance of weak supervision for a wide range of label models, classifiers, and datasets. Using less weakly-labeled data improves the accuracy of weak supervision pipelines by up to 19% (absolute) on benchmark tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Theoretical Analysis of Weak-to-Strong GeneralizationHunter Lang, David A. Sontag, Aravindan VijayaraghavanNeurIPS 2024 · 被引用 59 次
- Neighborhood-Regularized Self-Training for Learning with Few LabelsRan Xu, Yue Yu, Hejie Cui, Xuan Kan 等AAAI 2023 · 被引用 29 次
- SLaM: Student-Label Mixing for Distillation with Unlabeled ExamplesVasilis Kontonis, Fotis Iliopoulos, Khoa Trinh, Cenk Baykal 等NeurIPS 2023 · 被引用 10 次
- Characterizing the Impacts of Semi-supervised Learning for Weak SupervisionJeffrey Li, Jieyu Zhang, Ludwig Schmidt, Alexander J. RatnerNeurIPS 2023 · 被引用 9 次
- WISER: Weak Supervision and Supervised Representation Learning to Improve Drug Response Prediction in CancerKumar Shubham, Aishwarya Jayagopal, Syed Mohammed Danish, Prathosh A. P. 等ICML 2024 · 被引用 8 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 被引用 182 次
相关 Paper
- FastClass: A Time-Efficient Approach to Weakly-Supervised Text ClassificationTingyu Xia, Yue Wang, Yuan Tian, Yi ChangEMNLP 2022 · 被引用 1 次
- Weaker Than You Think: A Critical Look at Weakly Supervised LearningDawei Zhu, Xiaoyu Shen, Marius Mosbach, Andreas Stephan 等ACL 2023 · 被引用 13 次
- Training Data Subset Selection for Regression with Controlled Generalization ErrorDurga Sivasubramanian, Rishabh K. Iyer, Ganesh Ramakrishnan, Abir DeICML 2021 · 被引用 25 次
- The Perils of Learning From Unlabeled Data: Backdoor Attacks on Semi-supervised LearningVirat Shejwalkar, Lingjuan Lyu, Amir HoumansadrICCV 2023 · 被引用 15 次
- Boosting Few-Shot Visual Learning With Self-SupervisionSpyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez 等ICCV 2019 · 被引用 445 次
