RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Zijing Zhang, Ziyang Chen, Mingxiao Li, Zhaopeng Tu, Xiaolong Li
摘要
We demonstrate the effectiveness of RLVMR on two challenging long-horizon benchmarks, ALF-World and ScienceWorld. Our experiments show that RLVMR achieves new state-of-the-art results across all settings. Notably, on the hardest unseen task split (L2), our 7B model achieves an 83.6% success rate, and surpasses the performance of the much larger models. In-depth analysis reveals that these gains are driven by a tangible improvement in reasoning quality: RLVMR-trained agents exhibit significant reductions in repetitive and invalid actions. This confirms that by rewarding the process of good reasoning, we create agents that are not only more successful but also more robust, efficient, and generalizable. In summary, our contributions are as follows: 1. We identify and formulate the inefficient exploration problem in long-horizon agents, showing how optimizing for final outcomes alone reinforces flawed reasoning and leads to brittle policies that fail to generalize. 2. We propose RLVMR, a novel RL framework that provides dense, process-level supervision by rewarding verifiable meta-reasoning behaviors (e.g., planning, exploration, reflection) using lightweight, programmatic rules. 3. We achieve state-of-the-art performance on the challenging ALFWorld and ScienceWorld benchmarks, with significant improvements in generalization to unseen tasks. 4. We provide in-depth analysis confirming that RLVMR's gains stem directly from improved reasoning quality, evidenced by measurable reductions in repetitive actions and enhanced error recovery, thereby improving both agent robustness and efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningYulei Qin, Xiaoyu Tan, Zhengbao He, Gang Li 等ICLR 2026 · 被引用 9 次
- SPARK: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic LearningJinyang Wu, Shuo Yang, Yuhao Shen, Shuai Zhang 等ACL 2026 · 被引用 9 次
- Milestone-Guided Policy Learning for Long-Horizon Language AgentsZixuan Wang, Yuchen Yan, Hongxing Li, Teng Pan 等ICML 2026 · 被引用 8 次
- RoboAgent: Chaining Basic Capabilities for Embodied Task PlanningPeiran Xu, Jiaqi Zheng, Yadong MuCVPR 2026 · 被引用 6 次
- Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI AutomationZehao Deng, Tianjie Ju, Zheng Wu, Zhuosheng Zhang 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
相关 Paper
- GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-Based VLM Agent TrainingTong Wei, Yijun Yang, Junliang Xing, Yuanchun Shi 等ICCV 2025 · 被引用 1 次
- RLVR-World: Training World Models with Reinforcement LearningJialong Wu, Shaofeng Yin, Ningya Feng, Mingsheng LongNeurIPS 2025 · 被引用 52 次
- ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State ReflectionJeonghye Kim, Sojeong Rhee, Minbeom Kim, Dohyung Kim 等EMNLP 2025 · 被引用 2 次
- Promoting Efficient Reasoning with Verifiable Stepwise RewardChuhuai Yue, Chengqi Dong, Yinan Gao, Hang He 等AAAI 2026 · 被引用 19 次
- Dyna-Mind: Learning to Simulate from Experience for Better AI AgentsXiao Yu, Baolin Peng, Michel Galley, Hao Cheng 等ICLR 2026 · 被引用 8 次
