Reinforcement Learning Guided Semi-Supervised Learning
Marzi Heidari, Hanping Zhang, Yuhong Guo
摘要
In recent years, semi-supervised learning (SSL) has gained significant attention due to its ability to leverage both labeled and unlabeled data to improve model performance, especially when labeled data is scarce. However, most current SSL methods rely on heuristics or predefined rules for generating pseudo-labels and leveraging unlabeled data. They are limited to exploiting loss functions and regularization methods within the standard norm. In this paper, we propose a novel Reinforcement Learning (RL) Guided SSL method, RLGSSL, that formulates SSL as a one-armed bandit problem and deploys an innovative RL loss based on weighted reward to adaptively guide the learning process of the prediction model. RLGSSL incorporates a carefully designed reward function that balances the use of labeled and unlabeled data to enhance generalization performance. A semi-supervised teacher-student framework is further deployed to increase the learning stability. We demonstrate the effectiveness of RLGSSL through extensive experiments on several benchmark datasets and show that our approach achieves consistent superior performance compared to state-of-the-art SSL methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM ReasoningShenzhi Yang, Guangcheng Zhu, Haobo Wang, Xing Zheng 等ICLR 2026 · 被引用 6 次
- PCSR: Pseudo-label Consistency-Guided Sample Refinement for Noisy Correspondence LearningZhuoyao Liu, Yang Liu, Wentao Feng, Shudong HuangAAAI 2026
- Learning to Clean: Reinforcement Learning for Noisy Label CorrectionMarzi Heidari, Hanping Zhang, Yuhong GuoNeurIPS 2025
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu 等NeurIPS 2021 · 被引用 1,389 次
- ReMixMatch: Semi-Supervised Learning with Distribution Matching and Augmentation AnchoringDavid Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin 等ICLR 2020 · 被引用 469 次
相关 Paper
- Mitigating Data Scarcity in Supervised Machine Learning Through Reinforcement Learning Guided Data GenerationChengliang Chai, Kaisen Jin, Nan Tang, Ju Fan 等ICDE 2024 · 被引用 7 次
- LaSSL: Label-Guided Self-Training for Semi-supervised LearningZhen Zhao, Luping Zhou, Lei Wang, Yinghuan Shi 等AAAI 2022 · 被引用 51 次
- Bi-Level Optimization for Semi-Supervised Learning with Pseudo-LabelingMarzi Heidari, Yuhong GuoAAAI 2025 · 被引用 1 次
- PULNS: Positive-Unlabeled Learning with Effective Negative Sample SelectorChuan Luo, Pu Zhao, Chen Chen, Bo Qiao 等AAAI 2021 · 被引用 48 次
- SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement LearningJongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee 等ICLR 2022 · 被引用 116 次
