Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
Yoonjeon Kim, Doohyuk Jang, Eunho Yang
Abstract
Recent research on reasoning models explores the meta-awareness of language models, including their ability to determine optimal thinking duration, recognize knowledge boundaries, and structure concept-level thinking. While current large reasoning models depend solely on answer-based verification, we show that adding meta-awareness objectives leads to significant performance gains over models without such meta-knowledge. MAPR (Meta-Awareness via Predictive Reward) utilizes a self-generated task of predicting rollout statistics -specifically length, pass-rate, and concepts used -allowing for verification against the actual statistics. Furthermore, by leveraging this self-predictive capability, the model can regulate its reasoning behavior by i) filtering out trivial or unsolvable prompts, ii) reducing lengthy generations that tend to be incorrect, and iii) generating hints relevant to the problem. The results are inspiring: MAPR yields significant improvements in both accuracy and training efficiency on various reasoning benchmarks. More specifically, our method can speed up GRPO training by over 1.28× to reach the same performance, and achieve 83.18% gain in accuracy on AIME25, and a 13.04% average gain over six mathematics benchmarks. The code is publicly available at https://github.com /akatigre/MAPR-RL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a55cf46-1729-4d99-b770-c776ea39d164Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama et al.ICML 2024 · 270 citations
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu et al.ICLR 2026 · 250 citations
Related papers
- Step-GRPO: Enhancing Reasoning Quality and Efficiency via Structured PRM-Based Reinforcement LearningWeijie Li, Jin Wang, Liang-Chih Yu, Xuejie ZhangAAAI 2026 · 1 citation
- DRPO: Efficient Reasoning via Decoupled Reward Policy OptimizationGang Li, Yan Chen, Ming Lin, Tianbao YangICLR 2026 · 19 citations
- G²RPO-A: Guided Group Relative Policy Optimization with Adaptive GuidanceYongxin Guo, Wenbo Deng, Zhenglin Cheng, Xiaoying TangACL 2026 · 9 citations
- SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model ReasoningChenzhi Hu, Qinzhe Hu, Yuhang Xu, Junyi Chen et al.ICML 2026 · 2 citations
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective RolloutsHaizhong Zheng, Yang Zhou, Brian R. Bartoldson, Bhavya Kailkhura et al.NeurIPS 2025 · 125 citations
