Reinforcement Learning Guided Semi-Supervised Learning
Marzi Heidari, Hanping Zhang, Yuhong Guo
Abstract
In recent years, semi-supervised learning (SSL) has gained significant attention due to its ability to leverage both labeled and unlabeled data to improve model performance, especially when labeled data is scarce. However, most current SSL methods rely on heuristics or predefined rules for generating pseudo-labels and leveraging unlabeled data. They are limited to exploiting loss functions and regularization methods within the standard norm. In this paper, we propose a novel Reinforcement Learning (RL) Guided SSL method, RLGSSL, that formulates SSL as a one-armed bandit problem and deploys an innovative RL loss based on weighted reward to adaptively guide the learning process of the prediction model. RLGSSL incorporates a carefully designed reward function that balances the use of labeled and unlabeled data to enhance generalization performance. A semi-supervised teacher-student framework is further deployed to increase the learning stability. We demonstrate the effectiveness of RLGSSL through extensive experiments on several benchmark datasets and show that our approach achieves consistent superior performance compared to state-of-the-art SSL methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00ea997f-aec5-4a68-ab70-5f34654288a4Cited by top-tier papers3
- TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM ReasoningShenzhi Yang, Guangcheng Zhu, Haobo Wang, Xing Zheng et al.ICLR 2026 · 6 citations
- PCSR: Pseudo-label Consistency-Guided Sample Refinement for Noisy Correspondence LearningZhuoyao Liu, Yang Liu, Wentao Feng, Shudong HuangAAAI 2026
- Learning to Clean: Reinforcement Learning for Noisy Label CorrectionMarzi Heidari, Hanping Zhang, Yuhong GuoNeurIPS 2025
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu et al.NeurIPS 2021 · 1,389 citations
- ReMixMatch: Semi-Supervised Learning with Distribution Matching and Augmentation AnchoringDavid Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin et al.ICLR 2020 · 469 citations
Related papers
- Mitigating Data Scarcity in Supervised Machine Learning Through Reinforcement Learning Guided Data GenerationChengliang Chai, Kaisen Jin, Nan Tang, Ju Fan et al.ICDE 2024 · 7 citations
- LaSSL: Label-Guided Self-Training for Semi-supervised LearningZhen Zhao, Luping Zhou, Lei Wang, Yinghuan Shi et al.AAAI 2022 · 51 citations
- Bi-Level Optimization for Semi-Supervised Learning with Pseudo-LabelingMarzi Heidari, Yuhong GuoAAAI 2025 · 1 citation
- PULNS: Positive-Unlabeled Learning with Effective Negative Sample SelectorChuan Luo, Pu Zhao, Chen Chen, Bo Qiao et al.AAAI 2021 · 48 citations
- SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement LearningJongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee et al.ICLR 2022 · 116 citations
