Lune

KDD2026顶会

PROBE: Dense Process Rewards with Observation Evidence for Tool-Augmented Visual Reasoning

Zongsheng Cao, Anran Liu, Jun Xie, Feng Chen, Lang Chen, Zigan Wang

2026年份
1顶会引用

摘要

Tool-augmented reasoning offers a promising paradigm for improving multimodal large language models (MLLMs) by offloading perception to external tools. Yet existing methods leave three trajectory-level failures unsupervised: incomplete process supervision (sub-goals are skipped), tool hallucination (cited tools are not actually used), and observation hallucination (cited tool outputs are misquoted in the reasoning chain). We present PROBE, a rubric-centric framework in which a single set of atomic visual rubrics drives rubric-guided cold-start trajectory selection, a rubric-conditioned per-step tool reward, and the trajectory-level Visual Rubric Reward. On this rubric backbone, PROBE-GRPO further introduces Tool Evidence Chain Reward (graph-based tool-output utilisation) and Observation Faithfulness Reward (entailment-based content fidelity), and an Adaptive Learning scheme generalises to unseen tools via identifier obfuscation and description diversification. PROBE-7B surpasses proprietary models on multiple structured-reasoning benchmarks and exhibits self-adaptive tool-use behaviours under novel tool definitions.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖