AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
Mingyang Song, Haoyu Sun, Jiawei Gu, Linjie Li, Ranjay Krishna, Yu Cheng
摘要
While augmenting Multimodal Large Language Models (MLLMs) with tools is a promising direction, current approaches face critical limitations. They often rely on single, atomic tools, failing to address the challenges of multi-turn planning, and they do not equip models with the ability to select effective tool combinations for complex tasks. To overcome these limitations, we introduce AdaReasoner, a framework that teaches models to perform dynamic tool orchestration for iterative visual reasoning. Our paradigm is designed to support a broad spectrum of tools, including computationally intensive, expert-model-based services. It features a comprehensive design that includes a new data curation methodology and a tailored Tool GRPO algorithm to optimize multi-turn tool-calling trajectories, which yields state-of-the-art models that achieve substantial gains over their baselines (+38.7% average on 7B) and reach near-perfect accuracy on complex benchmarks like Visual Spatial Planning (97.6%). This performance surpasses leading proprietary systems such as GPT-5 and Claude Sonnet 4, demonstrating that our approach can effectively overcome scale-based limitations by augmenting smaller models with powerful tool-use capabilities. Critically, we find that AdaReasoner develops emergent, self-adaptive behaviors: it learns to autonomously adopt beneficial tools, discard irrelevant ones, and modulate its usage frequency. This ability to curate its own optimal problem-solving strategies represents a significant step toward building more robust, scalable, and reliable reasoning agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth 等NeurIPS 2024 · 被引用 373 次
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao 等ICLR 2026 · 被引用 321 次
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsShengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma 等CVPR 2024 · 被引用 111 次
相关 Paper
- AutoTool: Dynamic Tool Selection and Integration for Agentic ReasoningJiaru Zou, Ling Yang, Yunzhe Qi, Sirui Chen 等ICML 2026 · 被引用 4 次
- OctoTools: A Multi-Agent Framework with Extensible Tools for Complex ReasoningPan Lu, Bowen Chen, Sheng Liu, Rahul Thapa 等ACL 2026 · 被引用 2 次
- Learning to Inference Adaptively for Multimodal Large Language ModelsZhuoyan Xu, Khoi Duc Nguyen, Preeti Mukherjee, Saurabh Bagchi 等ICCV 2025 · 被引用 4 次
- PROBE: Dense Process Rewards with Observation Evidence for Tool-Augmented Visual ReasoningZongsheng Cao, Anran Liu, Jun Xie, Feng Chen 等KDD 2026
- GThinker: Towards General Multimodal Reasoning via Cue-Guided RethinkingYufei Zhan, Ziheng Wu, Yousong Zhu, Rongkun Xue 等CVPR 2026 · 被引用 14 次
