PROBE: Dense Process Rewards with Observation Evidence for Tool-Augmented Visual Reasoning
Zongsheng Cao, Anran Liu, Jun Xie, Feng Chen, Lang Chen, Zigan Wang
摘要
Tool-augmented reasoning offers a promising paradigm for improving multimodal large language models (MLLMs) by offloading perception to external tools. Yet existing methods leave three trajectory-level failures unsupervised: incomplete process supervision (sub-goals are skipped), tool hallucination (cited tools are not actually used), and observation hallucination (cited tool outputs are misquoted in the reasoning chain). We present PROBE, a rubric-centric framework in which a single set of atomic visual rubrics drives rubric-guided cold-start trajectory selection, a rubric-conditioned per-step tool reward, and the trajectory-level Visual Rubric Reward. On this rubric backbone, PROBE-GRPO further introduces Tool Evidence Chain Reward (graph-based tool-output utilisation) and Observation Faithfulness Reward (entailment-based content fidelity), and an Adaptive Learning scheme generalises to unseen tools via identifier obfuscation and description diversification. PROBE-7B surpasses proprietary models on multiple structured-reasoning benchmarks and exhibits self-adaptive tool-use behaviours under novel tool definitions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model ReasoningYukun Chen, Jiaming Li, Longze Chen, Ze Gong 等ICML 2026 · 被引用 5 次
- Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain ReasoningBaolong Bi, Shenghua Liu, Yiwei Wang, Siqian Tong 等ICML 2026 · 被引用 19 次
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented AgentsWonjoong Kim, Sangwu Park, Yeonjun In, Sein Kim 等ICML 2026 · 被引用 15 次
- On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMsRosie Zhao, Anshul Shah, Xiaoyu Zhu, Xinke Deng 等ICML 2026 · 被引用 10 次
- Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning ModelsGengwei Zhang, Jie Peng, Zhen Tan, Mufan Qiu 等CVPR 2026 · 被引用 1 次
