Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
Zizhuo Zhang, Jianing Zhu, Xinmu Ge, Zihua Zhao, Zhanke Zhou, Xuan Li, Xiao Feng, Jiangchao Yao, Bo Han
Abstract
While reinforcement learning with verifiable rewards (RLVR) is effective to improve the reasoning ability of large language models (LLMs), its reliance on human-annotated labels leads to the scaling up dilemma, especially for complex tasks. Recent self-rewarding methods investigate a label-free alternative to unlock the reasoning capabilities of LLMs, yet they frequently encounter the non-negligible training collapse issue, as the single-view supervision signal easily forms the self-consistent illusion, yielding the reward hacking. Inspired by the success of self-supervised learning, we propose Co-rewarding, a novel self-supervised RL framework that improves training stability by seeking complementary supervision from another views. Specifically, we instantiate Co-rewarding in two ways: (1) Co-rewarding-I is a data-side instantiation that derives reward signals from contrastive agreement across semantically analogous questions; and (2) Co-rewarding-II is a model-side instantiation that maintains a slowly-updated reference teacher with pseudo labels to realize self-distillation. Intuitively, such instantiations introduce different levels of discrepancy to increase the difficulty of training collapse on trivial reasoning solutions. We also explore their orthogonally combined version to further boost the performance. Empirically, Co-rewarding exhibits stable training across various setups, and outperforms other self-rewarding baselines by improvements on average on multiple mathematical reasoning benchmarks, especially by on Llama-3.2-3B-Instruct. Notably, Co-rewarding reaches or even surpasses RLVR with ground-truth (GT) label in several cases, such as a Pass@ of on GSM8K with Qwen3-8B-Base remarkably higher than GT. Our code is released at https://github.com/tmlr-group/Co-rewarding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f7530cf-a66a-48fe-8c84-e4c07449f052Cited by top-tier papers12
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement LearningYuhong Liu, Beichen Zhang, Yuhang Zang, Yuhang Cao et al.CVPR 2026 · 43 citations
- Landscape of Thoughts: Visualizing the Reasoning Process of Large Language ModelsZhanke Zhou, Zhaocheng Zhu, Xuan Li, Mikhail Galkin et al.ICLR 2026 · 28 citations
- Towards Understanding Valuable Preference Data for Large Language Model AlignmentZizhuo Zhang, Qizhou Wang, Shanshan Ye, Jianing Zhu et al.ICLR 2026 · 6 citations
- Reference-guided Policy Optimization for Molecular Optimization via LLM ReasoningXuan Li, Zhanke Zhou, Zongze Li, Jiangchao Yao et al.ICLR 2026 · 5 citations
- Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement LearningWeiqin Wang, Yile Wang, Kehao Chen, Hui HuangACL 2026 · 5 citations
Builds on45
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- CoAct: Co-Active LLM Preference Learning with Human-AI SynergyRuiyao Xu, Mihir Parmar, Tiankai Yang, Zhengyu Hu et al.ACL 2026 · 1 citation
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to DeliberationZhiwei Zhang, Xiaomin Li, Yudi Lin, Hui Liu et al.ICLR 2026 · 13 citations
- CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-EvolutionTeng Pan, Yuchen Yan, Zixuan Wang, Ruiqing Zhang et al.ACL 2026 · 5 citations
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui et al.NeurIPS 2025 · 20 citations
