Rethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards
Xinyu Tang, Yuliang Zhan, Zhixun Li, Xin Zhao, Zhenduo Zhang, Zujie Wen, Zhiqiang Zhang, Jun Zhou
Abstract
Large reasoning models (LRMs) are typically trained using reinforcement learning with verifiable reward (RLVR) to enhance their reasoning abilities. In this paradigm, policies are updated using both positive and negative self-generated rollouts, which correspond to distinct sample polarities. In this paper, we provide a systematic investigation into how these sample polarities affect RLVR training dynamics and behaviors. We find that positive samples sharpen existing correct reasoning patterns, while negative samples encourage exploration of new reasoning paths. We further explore how adjusting the advantage values of positive and negative samples at both the sample level and the token level affects RLVR training. Based on these insights, we propose an Adaptive and Asymmetric token-level Advantage shaping method for Policy Optimization, namely A3PO, that more precisely allocates advantage signals to key tokens across different polarities. Experiments across five reasoning benchmarks demonstrate the effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Reinforced Informativeness Optimization for Long-Form Retrieval-Augmented GenerationYuhao Wang, Ruiyang Ren, Yucheng Wang, Xin Zhao et al.ACL 2026 · 4 citations
- Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM ReasoningQin-Wen Luo, Sheng Ren, Xiang Chen, Rui Liu et al.KDD 2026 · 4 citations
- CloDS: Visual-Only Unsupervised Cloth Dynamics Learning in Unknown ConditionsYu-Liang Zhan, Jian Li, Wenbing Huang, Yang Liu et al.ICLR 2026
Builds on1
Related papers
- Experience Augmented Policy Optimization for LLM ReasoningJinda Lu, Kexin Huang, Junkang Wu, Shuo Yang et al.ICML 2026 · 2 citations
- FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable ReasoningYuyang Ding, Chi Zhang, Juntao Li, Haibin Lin et al.ICLR 2026 · 7 citations
- ExGRPO: Learning to Reason from ExperienceRunzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu et al.ICLR 2026 · 51 citations
- One-Way Policy Optimization for Self-Evolving LLMsShuo Yang, Jinda Lu, Kexin Huang, Chiyu Ma et al.ICML 2026 · 2 citations
- EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-ForgetLiang Chen, Xueting Han, Qizhou Wang, Bo Han et al.ICLR 2026 · 16 citations
