RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Zijing Zhang, Ziyang Chen, Mingxiao Li, Zhaopeng Tu, Xiaolong Li
Abstract
We demonstrate the effectiveness of RLVMR on two challenging long-horizon benchmarks, ALF-World and ScienceWorld. Our experiments show that RLVMR achieves new state-of-the-art results across all settings. Notably, on the hardest unseen task split (L2), our 7B model achieves an 83.6% success rate, and surpasses the performance of the much larger models. In-depth analysis reveals that these gains are driven by a tangible improvement in reasoning quality: RLVMR-trained agents exhibit significant reductions in repetitive and invalid actions. This confirms that by rewarding the process of good reasoning, we create agents that are not only more successful but also more robust, efficient, and generalizable. In summary, our contributions are as follows: 1. We identify and formulate the inefficient exploration problem in long-horizon agents, showing how optimizing for final outcomes alone reinforces flawed reasoning and leads to brittle policies that fail to generalize. 2. We propose RLVMR, a novel RL framework that provides dense, process-level supervision by rewarding verifiable meta-reasoning behaviors (e.g., planning, exploration, reflection) using lightweight, programmatic rules. 3. We achieve state-of-the-art performance on the challenging ALFWorld and ScienceWorld benchmarks, with significant improvements in generalization to unseen tasks. 4. We provide in-depth analysis confirming that RLVMR's gains stem directly from improved reasoning quality, evidenced by measurable reductions in repetitive actions and enhanced error recovery, thereby improving both agent robustness and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9f5b3e7-ba15-4053-8467-a5af074c0bc9Cited by top-tier papers11
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningYulei Qin, Xiaoyu Tan, Zhengbao He, Gang Li et al.ICLR 2026 · 9 citations
- SPARK: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic LearningJinyang Wu, Shuo Yang, Yuhao Shen, Shuai Zhang et al.ACL 2026 · 9 citations
- Milestone-Guided Policy Learning for Long-Horizon Language AgentsZixuan Wang, Yuchen Yan, Hongxing Li, Teng Pan et al.ICML 2026 · 8 citations
- RoboAgent: Chaining Basic Capabilities for Embodied Task PlanningPeiran Xu, Jiaqi Zheng, Yadong MuCVPR 2026 · 6 citations
- Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI AutomationZehao Deng, Tianjie Ju, Zheng Wu, Zhuosheng Zhang et al.CVPR 2026 · 1 citation
Builds on18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
Related papers
- GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-Based VLM Agent TrainingTong Wei, Yijun Yang, Junliang Xing, Yuanchun Shi et al.ICCV 2025 · 1 citation
- RLVR-World: Training World Models with Reinforcement LearningJialong Wu, Shaofeng Yin, Ningya Feng, Mingsheng LongNeurIPS 2025 · 52 citations
- ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State ReflectionJeonghye Kim, Sojeong Rhee, Minbeom Kim, Dohyung Kim et al.EMNLP 2025 · 2 citations
- Promoting Efficient Reasoning with Verifiable Stepwise RewardChuhuai Yue, Chengqi Dong, Yinan Gao, Hang He et al.AAAI 2026 · 19 citations
- Dyna-Mind: Learning to Simulate from Experience for Better AI AgentsXiao Yu, Baolin Peng, Michel Galley, Hao Cheng et al.ICLR 2026 · 8 citations
