Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable Rewards
Reinhard Heckel, Mahdi Soltanolkotabi, Christos Thrampoulidis
摘要
Reinforcement learning with verifiable rewards has driven recent advances in LLM post-training, in particular for reasoning. Policy optimization algorithms generate a number of responses for a given prompt and then effectively weight the corresponding gradients depending on the rewards. The most popular algorithms including GRPO, DAPO, and RLOO focus on ambiguous prompts, i.e., prompts with intermediate success probability, while downgrading gradients with very easy and very hard prompts. In this paper, we consider asymmetric prompt weightings that assign higher weights to prompts with low, or even zero, empirical success probability. We find that asymmetric weighting particularly benefits from-scratch RL (as in R1-Zero), where training traverses a wide accuracy range, and less so in post-SFT RL where the model already starts at high accuracy. We also provide theory that characterizes prompt weights which minimize the time needed to raise success probability from an initial level to a target accuracy under a fixed update budget. In low-success regimes, where informative responses are rare and response cost dominates, these optimal weights become asymmetric, upweighting low success probabilities and thereby accelerating effective-time convergence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 被引用 59 次
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage ShapingThanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Dac Lai 等ICLR 2026 · 被引用 55 次
- Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMsArash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee 等ACL 2024 · 被引用 20 次
相关 Paper
- Multi-Task GRPO: Reliable LLM Reasoning Across TasksShyam Sundhar Ramesh, Xiaotong Ji, Matthieu Zimmer, Sangwoong Yoon 等ICML 2026 · 被引用 8 次
- XRPO: Pushing the Limits of GRPO with Targeted Exploration and ExploitationUdbhav Bamba, Minghao Fang, Yifan Yu, Haizhong Zheng 等ICML 2026 · 被引用 17 次
- The Easy, the Hard, and the Learnable: Confidence and Difficulty-Adaptive Policy Optimization for LLM Reasoning(Andrew) Zhanke Zhou, Xiangyu Lu, Chentao Cao, Brando Miranda 等ICML 2026
- One-Way Policy Optimization for Self-Evolving LLMsShuo Yang, Jinda Lu, Kexin Huang, Chiyu Ma 等ICML 2026 · 被引用 2 次
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy EntropyHongze Tan, Zihan Wang, Jianfei Pan, Jinghao Lin 等ICML 2026 · 被引用 53 次
