Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?
Yun Qu, Qi Wang, Yixiu Mao, Vincent Tao Hu, Björn Ommer, Xiangyang Ji
Abstract
Recent advances have witnessed the effectiveness of reinforcement learning (RL) finetuning in enhancing the reasoning capabilities of large language models (LLMs). The optimization process often requires numerous iterations to achieve satisfactory performance, resulting in high computational costs due to the need for frequent prompt evaluations under intensive LLM interactions and repeated policy updates. Appropriate online prompt selection methods reduce iteration steps by prioritizing informative prompts during training, while the pipeline's reliance on exhaustive prompt evaluation and subset selection for optimization still incurs substantial computational overhead due to frequent LLM inference calls. Distinguished from these direct evaluate-then-select schemes, this work investigates iterative approximate evaluation for arbitrary prompts and introduces Model Predictive Prompt Selection (MoPPS), a Bayesian risk-predictive framework that online estimates prompt difficulty without requiring costly LLM interactions. Technically, MoPPS models each prompt's success rate as a latent variable, performs streaming Bayesian inference, and employs posterior sampling in a constructed multi-armed bandit machine, enabling efficient and adaptive prompt selection. Extensive experiments across mathematics, planning, and vision-based geometry tasks show that MoPPS reliably predicts prompt difficulty and accelerates training with significantly reduced LLM rollouts. Our code is available at https://github.com/thu-rllab/MoPPS . CCS Concepts • Computing methodologies → Artificial intelligence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2cda51d-963c-4451-942b-55828014285dCited by top-tier papers20
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage ShapingThanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Dac Lai et al.ICLR 2026 · 55 citations
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-ExpertsHeming Zou, Yunliang Zang, Wutong Xu, Yao Zhu et al.NeurIPS 2025 · 38 citations
- Diversity-Incentivized Exploration for Versatile ReasoningZican Hu, Shilin Zhang, Yafu Li, Jianhao Yan et al.ICLR 2026 · 32 citations
- Single-stream Policy OptimizationZhongwen Xu, Zihan DingICLR 2026 · 29 citations
- CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMsYongcheng Zeng, Zexu Sun, Bokai Ji, Erxue Min et al.ICLR 2026 · 18 citations
Builds on14
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang et al.NeurIPS 2025 · 533 citations
Related papers
- Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning ModelsYixiu Mao, Yun Qu, Qi Wang, Heming Zou et al.ICLR 2026 · 13 citations
- Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning ModelsYun Qu, Qi Wang, Yixiu Mao, Heming Zou et al.ICML 2026 · 7 citations
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective RolloutsHaizhong Zheng, Yang Zhou, Brian R. Bartoldson, Bhavya Kailkhura et al.NeurIPS 2025 · 125 citations
- Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RLHao Sun, Alihan Hüyük, Mihaela van der SchaarICLR 2024 · 48 citations
- Hyperband-based Bayesian Optimization for Black-box Prompt SelectionLennart Schneider, Martin Wistuba, Aaron Klein, Jacek Golebiowski et al.ICML 2025
