Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
Tianyi Men, Zhuoran Jin, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
摘要
As Multimodal Large Language Models (MLLMs) advance, multimodal agents show promise in real-world tasks like web navigation and embodied intelligence. However, due to limitations in a lack of external feedback, these agents struggle with self-correction and generalization. A promising approach is to use reward models as external feedback, but there is no clear on how to select reward models for agents. Thus, there is an urgent need to build a reward bench targeted at agents. To address these challenges, we propose Agent-RewardBench, a benchmark designed to evaluate reward modeling ability in MLLMs. The benchmark is characterized by three key features: (1) Multiple dimensions and real-world agent scenarios evaluation. It covers perception, planning, and safety with 7 scenarios; (2) Step-level reward evaluation. It allows for the assessment of agent capabilities at the individual steps of a task, providing a more granular view of performance during the planning process; and (3) Appropriately difficulty and high-quality. We carefully sample from 10 diverse models, difficulty control to maintain task challenges, and manual verification to ensure the integrity of the data. Experiments demonstrate that even state-of-the-art multimodal models show limited performance, highlighting the need for specialized training in agent reward modeling. Code is available at github.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- CUARewardBench: A Benchmark for Evaluating Reward Models for Computer-Using AgentsHaojia Lin, Xiaoyu Tan, Yulei Qin, Zihan Xu 等ICML 2026 · 被引用 15 次
- Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded VerificationMoises Andrade, Joonhyuk Cha, Brandon Ho, Vriksha Srihari 等ICLR 2026 · 被引用 12 次
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form PreferencesZhuoran Jin, Hongbang Yuan, Kejian Zhu, Jiachun Li 等ICLR 2026 · 被引用 11 次
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic ModelsZhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang 等CVPR 2026 · 被引用 6 次
- IF-RewardBench: Benchmarking Judge Models for Instruction-Following EvaluationBosi Wen, Yilin Niu, Cunxiang Wang, Xiaoying Ling 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper12
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- TravelPlanner: A Benchmark for Real-World Planning with Language AgentsJian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu 等ICML 2024 · 被引用 376 次
- OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task SynthesisQiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin 等ACL 2025 · 被引用 114 次
- Attacking Vision-Language Computer Agents via Pop-upsYanzhe Zhang, Tao Yu, Diyi YangACL 2025 · 被引用 99 次
相关 Paper
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsRui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao 等ICML 2025
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsChun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta Ying Cheng 等AAAI 2026 · 被引用 35 次
- Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward ModelingJiaxuan Wang, Yulan Hu, Wenjin Yang, Zheng Pan 等ACL 2026 · 被引用 1 次
- MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QAXianjie Wu, Xiaohang Xu, Tingyu Jiang, Jian Yang 等WWW 2026 · 被引用 3 次
- NavBench: Probing Multimodal Large Language Models for Embodied NavigationYanyuan Qiao, Haodong Hong, Wenqi Lyu, Dong An 等NeurIPS 2025 · 被引用 27 次
