PROBE: Dense Process Rewards with Observation Evidence for Tool-Augmented Visual Reasoning
Zongsheng Cao, Anran Liu, Jun Xie, Feng Chen, Lang Chen, Zigan Wang
Abstract
Tool-augmented reasoning offers a promising paradigm for improving multimodal large language models (MLLMs) by offloading perception to external tools. Yet existing methods leave three trajectory-level failures unsupervised: incomplete process supervision (sub-goals are skipped), tool hallucination (cited tools are not actually used), and observation hallucination (cited tool outputs are misquoted in the reasoning chain). We present PROBE, a rubric-centric framework in which a single set of atomic visual rubrics drives rubric-guided cold-start trajectory selection, a rubric-conditioned per-step tool reward, and the trajectory-level Visual Rubric Reward. On this rubric backbone, PROBE-GRPO further introduces Tool Evidence Chain Reward (graph-based tool-output utilisation) and Observation Faithfulness Reward (entailment-based content fidelity), and an Adaptive Learning scheme generalises to unseen tools via identifier obfuscation and description diversification. PROBE-7B surpasses proprietary models on multiple structured-reasoning benchmarks and exhibits self-adaptive tool-use behaviours under novel tool definitions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f2bea90e-0d2c-4f24-be39-39253935b785Cited by top-tier papers1
Ask how each one uses itRelated papers
- RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model ReasoningYukun Chen, Jiaming Li, Longze Chen, Ze Gong et al.ICML 2026 · 5 citations
- Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain ReasoningBaolong Bi, Shenghua Liu, Yiwei Wang, Siqian Tong et al.ICML 2026 · 19 citations
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented AgentsWonjoong Kim, Sangwu Park, Yeonjun In, Sein Kim et al.ICML 2026 · 15 citations
- On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMsRosie Zhao, Anshul Shah, Xiaoyu Zhu, Xinke Deng et al.ICML 2026 · 10 citations
- Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning ModelsGengwei Zhang, Jie Peng, Zhen Tan, Mufan Qiu et al.CVPR 2026 · 1 citation
