Lune

KDD2026Top-tier venue

PROBE: Dense Process Rewards with Observation Evidence for Tool-Augmented Visual Reasoning

Zongsheng Cao, Anran Liu, Jun Xie, Feng Chen, Lang Chen, Zigan Wang

2026Year
1Top-tier citations

Abstract

Tool-augmented reasoning offers a promising paradigm for improving multimodal large language models (MLLMs) by offloading perception to external tools. Yet existing methods leave three trajectory-level failures unsupervised: incomplete process supervision (sub-goals are skipped), tool hallucination (cited tools are not actually used), and observation hallucination (cited tool outputs are misquoted in the reasoning chain). We present PROBE, a rubric-centric framework in which a single set of atomic visual rubrics drives rubric-guided cold-start trajectory selection, a rubric-conditioned per-step tool reward, and the trajectory-level Visual Rubric Reward. On this rubric backbone, PROBE-GRPO further introduces Tool Evidence Chain Reward (graph-based tool-output utilisation) and Observation Faithfulness Reward (entailment-based content fidelity), and an Adaptive Learning scheme generalises to unseen tools via identifier obfuscation and description diversification. PROBE-7B surpasses proprietary models on multiple structured-reasoning benchmarks and exhibits self-adaptive tool-use behaviours under novel tool definitions.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get f2bea90e-0d2c-4f24-be39-39253935b785

Cited by top-tier papers1

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines