Discounted Beta–Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards
Haechan Kim, Soohyun Ryu, Gyouk Chu, Doohyuk Jang, Eunho Yang
摘要
Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective post-training paradigm for improving the reasoning capabilities of large language models. However, existing group-based RLVR methods often suffer from severe sample inefficiency. This inefficiency stems from reliance on point estimation of rewards from a small number of rollouts, leading to high estimation variance, variance collapse, and ineffective utilization of generated responses. In this work, we reformulate RLVR from a statistical estimation perspective by modeling rewards as samples drawn from a policy-induced distribution and casting advantage computation as the problem of estimating the reward distribution from finite data. Building on this view, we propose D iscounted B eta– B ernoulli ( DBB ) reward estimation, which leverages historical reward statistics for the non-stationary distribution. Although biased, the resulting estimator exhibits reduced and stable variance and theoretically avoids variance collapse. Under mild non-stationarity, it also achieves a lower mean squared error than standard point estimation, as we characterize analytically and verify empirically. Across six in-distribution and three out-of-distribution reasoning benchmarks, GRPO with DBB consistently outperforms naive GRPO and strong recent baselines, including the replay-based RePO and the variance-collapse-aware GRESO and DAPO. Relative to GRPO, it achieves average Acc@8 improvements of 3.43/2.32 points in-distribution and 10.05/8.34 points out-of-distribution on the 1.7B and 8B models, respectively, without additional computational cost or memory usage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective RolloutsHaizhong Zheng, Yang Zhou, Brian R. Bartoldson, Bhavya Kailkhura 等NeurIPS 2025 · 被引用 125 次
- Geometric-Mean Policy OptimizationYuzhong Zhao, Yue Liu, Junpeng Liu, Jingye Chen 等ICLR 2026 · 被引用 104 次
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage ShapingThanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Dac Lai 等ICLR 2026 · 被引用 55 次
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu 等ACL 2024 · 被引用 18 次
相关 Paper
- ExGRPO: Learning to Reason from ExperienceRunzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu 等ICLR 2026 · 被引用 51 次
- Stable and Efficient Single-Rollout RL for Multimodal ReasoningRui Liu, Dian Yu, Lei Ke, Haolin Liu 等CVPR 2026 · 被引用 13 次
- Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable RewardsGuanning Zeng, Zhaoyi Zhou, Daman Arora, Andrea ZanetteICML 2026 · 被引用 13 次
- Resource-Efficient Reinforcement for Reasoning Large Language Models via Dynamic One-Shot Policy RefinementYunjian Zhang, Sudong Wang, Yang Li, Peiran Xu 等ICML 2026 · 被引用 4 次
- Advantage Collapse in Group Relative Policy Optimization: Diagnosis and MitigationXixiang He, Qiyao Sun, Ao Cheng, Xingming Li 等ICML 2026
