Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
Christian Walder, Deep Karkhanis
摘要
Reinforcement Learning (RL) algorithms sample multiple n>1 solution attempts for each problem and reward them independently. This optimizes for pass@1 performance and prioritizes the strength of isolated samples at the expense of the diversity and collective utility of sets of samples. This under-utilizes the sampling capacity, limiting exploration and eventual improvement on harder examples. As a fix, we propose Pass-at-k Policy Optimization (PKPO), a transformation on the final rewards which leads to direct optimization of pass@k performance, thus optimizing for sets of samples that maximize reward when considered jointly. Our contribution is to derive novel low variance unbiased estimators for pass@k and its gradient, in both the binary and continuous reward settings. We show optimization with our estimators reduces to standard RL with rewards that have been jointly transformed by a stable and efficient transformation function. While previous efforts are restricted to k=n, ours is the first to enable robust optimization of pass@k for any arbitrary k<= n. Moreover, instead of trading off pass@1 performance for pass@k gains, our method allows annealing k during training, optimizing both metrics and often achieving strong pass@1 numbers alongside significant pass@k gains. We validate our reward transformations on toy experiments, which reveal the variance reducing properties of our formulations. We also include real-world examples using the open-source LLM, GEMMA-2. We find that our transformation effectively optimizes for the target k. Furthermore, higher k values enable solving more and harder problems, while annealing k boosts both the pass@1 and pass@k . Crucially, for challenging task sets where conventional pass@1 optimization stalls, our pass@k approach unblocks learning, likely due to better exploration by prioritizing joint utility over the utility of individual samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Maximum Likelihood Reinforcement LearningFahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song 等ICML 2026 · 被引用 18 次
- Representation-Based Exploration for Language Models: From Test-Time to Post-TrainingJens Tuyls, Dylan J Foster, Akshay Krishnamurthy, Jordan T. AshICLR 2026 · 被引用 18 次
- Differential Smoothing Mitigates Sharpening and Improves LLM ReasoningJingchu Gai, Guanning Zeng, Huaqing ZHANG, Aditi RaghunathanICML 2026 · 被引用 13 次
- Reinforcement Learning on Pre-Training DataSiheng Li, Kejiao Li, Zenan Xu, Guanhua Huang 等ACL 2026 · 被引用 11 次
- ProofOptimizer: Training Language Models to Simplify Proofs without Human DemonstrationsAlex Gu, Bartosz Piotrowski, Fabian Gloeckle, Kaiyu Yang 等ICLR 2026 · 被引用 11 次
它引用的顶会 Paper6
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese 等NeurIPS 2022 · 被引用 571 次
- Emergence of Exploration in Policy Gradient Reinforcement Learning via RetryingSoichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno 等ICML 2026 · 被引用 6 次
- Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language ModelsYinlam Chow, Guy Tennenholtz, Izzeddin Gur, Vincent Zhuang 等ICLR 2025 · 被引用 1 次
- Training Language Models to Self-Correct via Reinforcement LearningAviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su 等ICLR 2025
- BOND: Aligning LLMs with Best-of-N DistillationPier Giuseppe Sessa, Robert Dadashi-Tazehozi, Léonard Hussenot, Johan Ferret 等ICLR 2025
相关 Paper
- Group-Aware Reinforcement Learning for Output Diversity in Large Language ModelsOron Anschel, Alon Shoshan, Adam Botach, Shunit Haviv Hakimi 等EMNLP 2025 · 被引用 1 次
- Rewarding the Unlikely: Lifting GRPO Beyond Distribution SharpeningAndre Wang He, Daniel Fried, Sean WelleckEMNLP 2025 · 被引用 56 次
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMsXumeng Wen, Zihan Liu, Shun Zheng, Shengyu Ye 等ICLR 2026 · 被引用 279 次
- Risk-Sensitive Reinforcement Learning for Alleviating Exploration Dilemmas in Large Language ModelsYuhua Jiang, Jiawei Huang, Yufeng Yuan, Xin Mao 等ICLR 2026 · 被引用 8 次
- RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-TrainingTao Ren, Jinyang Jiang, Hui Yang, Wan Tian 等ICLR 2026 · 被引用 8 次
