No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
Thanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Dac Lai, Eunho Yang
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful framework for improving the reasoning abilities of Large Language Models (LLMs). However, current methods such as GRPO rely only on problems where the model responses to the same input differ in correctness, while ignoring those where all responses receive the same reward -so-called zero-variance prompts. In this work, we argue that such prompts are not useless but can, in fact, provide meaningful feedback for policy optimization. To this end, we introduce Reinforcement Learning with Zero-Variance Prompts (RL-ZVP), a novel algorithm that extract learning signals from zero-variance prompts. RL-ZVP directly rewards correctness and penalizes errors even without contrasting responses, modulating feedback with token-level characteristics to preserve informative, nuanced signals. Across six math reasoning benchmarks, RL-ZVP achieves significant improvements of up to 8.61 points in accuracy and 7.77 points in pass rate over GRPO, while consistently outperforming other baselines that filter out zero-variance prompts. These results highlight the untapped potential of learning from zero-variance prompts in RLVR. The project page is available at https: //bltnynk.github.io/publications/rl-zvp/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7021b32-564a-4c97-a7b7-58e533aa561bCited by top-tier papers6
- Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment PerspectiveDeyang Kong, Qi Guo, Xiangyu Xi, Wei Wang et al.AAAI 2026 · 5 citations
- Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable RewardsReinhard Heckel, Mahdi Soltanolkotabi, Christos ThrampoulidisICML 2026 · 2 citations
- On Effectiveness and Efficiency of Agentic Tool-calling and RL TrainingTong Liu, Cheng Qian, Matej Cief, Yuan He et al.ICML 2026 · 1 citation
- Advantage Collapse in Group Relative Policy Optimization: Diagnosis and MitigationXixiang He, Qiyao Sun, Ao Cheng, Xingming Li et al.ICML 2026
- Discounted Beta–Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable RewardsHaechan Kim, Soohyun Ryu, Gyouk Chu, Doohyuk Jang et al.ICML 2026
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng et al.NeurIPS 2025 · 592 citations
Related papers
- Risk-Sensitive Reinforcement Learning for Alleviating Exploration Dilemmas in Large Language ModelsYuhua Jiang, Jiawei Huang, Yufeng Yuan, Xin Mao et al.ICLR 2026 · 8 citations
- MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR AdvancementWeitao Jia, Jinghui Lu, Haiyang Yu, Siqi Wang et al.AAAI 2026 · 12 citations
- Random Policy Valuation is Enough for LLM Reasoning with Verifiable RewardsHaoran He, Yuxiao Ye, Qingpeng Cai, Chen Hu et al.ICLR 2026 · 9 citations
- ExGRPO: Learning to Reason from ExperienceRunzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu et al.ICLR 2026 · 51 citations
- Spurious Rewards: Rethinking Training Signals in RLVRRulin Shao, Stella Li, Rui Xin, Scott Geng et al.ICML 2026
