LLM-Guided Co-Training for Text Classification
Md Mezbaur Rahman, Cornelia Caragea
摘要
In this paper, we introduce a novel weighted co-training approach that is guided by Large Language Models (LLMs). Namely, in our co-training approach, we use LLM labels on unlabeled data as target labels and co-train two encoder-only based networks that train each other over multiple iterations: first, all samples are forwarded through each network and historical estimates of each network's confidence in the LLM label are recorded; second, a dynamic importance weight is derived for each sample according to each network's belief in the quality of the LLM label for that sample; finally, the two networks exchange importance weights with each other-each network back-propagates all samples weighted with the importance weights coming from its peer network and updates its own parameters. By strategically utilizing LLM-generated guidance, our approach significantly outperforms conventional SSL methods, particularly in settings with abundant unlabeled data. Empirical results show that it achieves state-of-the-art performance on 4 out of 5 benchmark datasets and ranks first among 14 compared methods according to the Friedman test. Our results highlight a new direction in semi-supervised learningwhere LLMs serve as knowledge amplifiers, enabling backbone co-training models to achieve state-of-the-art performance efficiently.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu 等NeurIPS 2021 · 被引用 1,389 次
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Dash: Semi-Supervised Learning with Dynamic ThresholdingYi Xu, Lei Shang, Jinxing Ye, Qi Qian 等ICML 2021 · 被引用 287 次
相关 Paper
- Reinforcement Learning Guided Semi-Supervised LearningMarzi Heidari, Hanping Zhang, Yuhong GuoNeurIPS 2024 · 被引用 6 次
- LaSSL: Label-Guided Self-Training for Semi-supervised LearningZhen Zhao, Luping Zhou, Lei Wang, Yinghuan Shi 等AAAI 2022 · 被引用 51 次
- All Labels Are Not Created Equal: Enhancing Semi-Supervision via Label Grouping and Co-TrainingIslam Nassar, Samitha Herath, Ehsan Abbasnejad, Wray L. Buntine 等CVPR 2021
- Co-training for Low Resource Scientific Natural Language InferenceMobashir Sadat, Cornelia CarageaACL 2024
- Bi-Level Optimization for Semi-Supervised Learning with Pseudo-LabelingMarzi Heidari, Yuhong GuoAAAI 2025 · 被引用 1 次
