Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark
Bingchen Miao, Yang Wu, Minghe Gao, Qifan Yu, Wendong Bu, Wenqiao Zhang, Yunfei Li, Siliang Tang, Tat-Seng Chua, Juncheng Li
摘要
The development of Generalist Virtual Agents (GVAs) has shown significant promise in autonomous task execution. However, current training paradigms face critical limitations, including reliance on outcome supervision and labor-intensive human annotations. To address these challenges, we propose Similar, a stepwise multi-dimensional generalist reward model, which offers fine-grained signals for agent training and can choose better actions for inferencetime scaling. Specifically, we begin by systematically defining five dimensions for evaluating agent actions. Building on this framework, we design an MCTS-P algorithm to automatically collect and annotate step-wise, five-dimensional agent execution data. Using this data, we train Similar with our crafted Triple-M strategy. Furthermore, we introduce the first benchmark in the virtual agent domain for step-wise, multi-dimensional reward model training and evaluation, named SRM. This benchmark consists of two components: SRMTrain, which serves as the training set for Similar, and SRMEval, a manually selected test set for evaluating the reward model. Experimental results demonstrate that Similar, through its step-wise, multi-dimensional assessment and synergistic gain, provides GVAs with effective intermediate signals during both training and inference-time scaling. The code is available in https://github.com/antgroup/Similar .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and ProgressZhiheng Xi, Chenyang Liao, Guanyu Li, Zhihao Zhang 等WWW 2026 · 被引用 19 次
- WebArbiter: A Generative Reasoning Process Reward Model for Web AgentsYao Zhang, Shijie Tang, Zeyu Li, Zhen Han 等ICLR 2026 · 被引用 2 次
- Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware ExplorationWeile Chen, Bingchen Miao, Qifan Yu, Wendong Bu 等CVPR 2026
- What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent CapabilitiesWendong Bu, Yang Wu, Qifan Yu, Minghe Gao 等ICML 2025
它引用的顶会 Paper21
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu 等ICLR 2024 · 被引用 637 次
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchDan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue 等NeurIPS 2024 · 被引用 527 次
相关 Paper
- Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal AgentsTianyi Men, Zhuoran Jin, Pengfei Cao, Yubo Chen 等ACL 2025 · 被引用 13 次
- Benchmarking Multimodal CoT Reward Model Stepwise by Visual ProgramMinghe Gao, Xuqi Liu, Zhongqi Yue, Yang Wu 等ICCV 2025
- SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy TasksWeikai Xu, Zhizheng Jiang, Yuxuan Liu, Pengzhi Gao 等ICLR 2026
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic ModelsZhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang 等CVPR 2026 · 被引用 6 次
- AgentStudio: A Toolkit for Building General Virtual AgentsLongtao Zheng, Zhiyuan Huang, Zhenghai Xue, Xinrun Wang 等ICLR 2025
