AgentRM: Enhancing Agent Generalization with Reward Modeling
Yu Xia, Jingru Fan, Weize Chen, Siyu Yan, Xin Cong, Zhong Zhang, Yaxi Lu, Yankai Lin, Zhiyuan Liu, Maosong Sun
摘要
Existing LLM-based agents have achieved strong performance on held-in tasks, but their generalizability to unseen tasks remains poor. Hence, some recent work focus on fine-tuning the policy model with more diverse tasks to improve the generalizability. In this work, we find that finetuning a reward model to guide the policy model is more robust than directly finetuning the policy model. Based on this finding, we propose AgentRM, a generalizable reward model, to guide the policy model for effective test-time search. We comprehensively investigate three approaches to construct the reward model, including explicit reward modeling, implicit reward modeling and LLM-as-a-judge. We then use AgentRM to guide the answer generation with Best-of-N sampling and step-level beam search. On four types of nine agent tasks, AgentRM enhances the base policy model by points on average, surpassing the top general agent by . Moreover, it demonstrates weak-to-strong generalization, yielding greater improvement of on LLaMA-3-70B policy model. As for the specializability, AgentRM can also boost a finetuned policy model and outperform the top specialized agent by on three held-in tasks. Further analysis verifies its effectiveness in test-time scaling. Codes will be released to facilitate the research in this area.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided ImprovementRuihan Yang, Fanghua Ye, Jian Li, Siyu Yuan 等NeurIPS 2025 · 被引用 21 次
- AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and ProgressZhiheng Xi, Chenyang Liao, Guanyu Li, Zhihao Zhang 等WWW 2026 · 被引用 19 次
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningYulei Qin, Xiaoyu Tan, Zhengbao He, Gang Li 等ICLR 2026 · 被引用 9 次
- Test-Time Deep Thinking to Explore Implicit RulesWentong Chen, Xin Cong, Zhong Zhang, Yaxi Lu 等KDD 2026
它引用的顶会 Paper13
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Large Language Models Can Self-ImproveJiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu 等EMNLP 2023 · 被引用 184 次
相关 Paper
- EQA-RM: A Generative Embodied Reward Model with Test-time ScalingYuhang Chen, Zhen Tan, Tianlong ChenEMNLP 2025 · 被引用 1 次
- HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented GenerationPeilin Wu, Mian Zhang, Kun Wan, Wentian Zhao 等ICLR 2026 · 被引用 13 次
- Pre-Trained Policy Discriminators are General Reward ModelsShihan Dou, Shichun Liu, Yuming Yang, Yicheng Zou 等NeurIPS 2025 · 被引用 13 次
- Empowering LLM Tool Invocation with Tool-call Reward ModelDa Ma, Ziyue Yang, Hongshen Xu, Haotian Fang 等ICLR 2026
- RM-R1: Reward Modeling as ReasoningXiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin 等ICLR 2026 · 被引用 147 次
