SemiReward: A General Reward Model for Semi-supervised Learning
Siyuan Li, Weiyang Jin, Zedong Wang, Fang Wu, Zicheng Liu, Cheng Tan, Stan Z. Li
摘要
Semi-supervised learning (SSL) has witnessed great progress with various improvements in the self-training framework with pseudo labeling. The main challenge is how to distinguish high-quality pseudo labels against the confirmation bias. However, existing pseudo-label selection strategies are limited to pre-defined schemes or complex hand-crafted policies specially designed for classification, failing to achieve high-quality labels, fast convergence, and task versatility simultaneously. To these ends, we propose a Semi-supervised Reward framework (SemiReward) that predicts reward scores to evaluate and filter out high-quality pseudo labels, which is pluggable to mainstream SSL methods in wide task types and scenarios. To mitigate confirmation bias, SemiReward is trained online in two stages with a generator model and subsampling strategy. With classification and regression tasks on 13 standard SSL benchmarks across three modalities, extensive experiments verify that SemiReward achieves significant performance gains and faster convergence speeds upon Pseudo Label, FlexMatch, and Free/-SoftMatch. Code and models are available at https://github.com/Westl ake-AI/SemiReward .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Harnessing Hard Mixed Samples with Decoupled RegularizerZicheng Liu, Siyuan Li, Ge Wang, Lirong Wu 等NeurIPS 2023 · 被引用 28 次
- Adversarial AutoMixupHuafeng Qin, Xin Jin, Yun Jiang, Mounîm A. El-Yacoubi 等ICLR 2024 · 被引用 19 次
- OmniGaze: Reward-inspired Generalizable Gaze Estimation in the WildHongyu Qu, Jianan Wei, Xiangbo Shu, Yazhou Yao 等NeurIPS 2025 · 被引用 15 次
- Instructor-inspired Machine Learning for Robust Molecular Property PredictionFang Wu, Shuting Jin, Siyuan Li, Stan Z. LiNeurIPS 2024 · 被引用 14 次
- When Confidence Fails: Revisiting Pseudo-Label Selection in Semi-Supervised Semantic SegmentationPan Liu, Jinshi LiuICCV 2025 · 被引用 9 次
它引用的顶会 Paper33
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
相关 Paper
- RegMixMatch: Optimizing Mixup Utilization in Semi-Supervised LearningHaorong Han, Jidong Yuan, Chixuan Wei, Zhongyang YuAAAI 2025 · 被引用 7 次
- Rethinking Confidence Scores and Thresholds in Pseudolabeling-based SSLHarit Vishwakarma, Yi Chen, Satya Sai Srinath Namburi GNVV, Sui Jiet Tay 等ICML 2025
- HyperMatch: Noise-Tolerant Semi-Supervised Learning via Relaxed Contrastive ConstraintBeitong Zhou, Jing Lu, Kerui Liu, Yunlu Xu 等CVPR 2023
- FreeMatch: Self-adaptive Thresholding for Semi-supervised LearningYidong Wang, Hao Chen, Qiang Heng, Wenxin Hou 等ICLR 2023 · 被引用 139 次
- SoftMatch: Addressing the Quantity-Quality Tradeoff in Semi-supervised LearningHao Chen, Ran Tao, Yue Fan, Yidong Wang 等ICLR 2023
