Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents
Dong-Hee Kim, Reuben Tan, DONGHYUN KIM
摘要
Visual agents employ external visual tools within visual chains of thought to incorporate fine-grained evidence. While prior work has mainly studied these tools in visual search tasks, their role in more complex visual reasoning remains underexplored. In this paper, we move beyond simple visual search tasks to investigate more challenging tasks, including 3D spatial reasoning and medical visual question answering, where agents must integrate tool-acquired local evidence with the global context. We identify a tool-use collapse phenomenon: models progressively stop using tools while still achieving higher task accuracy. Moreover, we observe a clear asymmetry: (i) completely eliminating tool use degrades performance, whereas (ii) incentivizing tool use yields only marginal gains despite substantially increasing usage. We find that vanilla training and tool-use encouragement both reduce rollout diversity, explaining why higher tool use does not yield stronger reasoning performance. Motivated by these findings, we add an entropy regularization term to encourage diverse rollout exploration, achieving the best performance despite gradually declining tool usage. Overall, our findings suggest a training-time view of tools as scaffolding, where broader exploration over language generation and visual tool invocation improves reasoning despite tool-use collapse. Project page: https://scaffolded-exploration.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth 等NeurIPS 2024 · 被引用 373 次
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao 等ICLR 2026 · 被引用 321 次
- VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool UseMingyuan Wu, Jingcheng Yang, Jize Jiang, Meitang Li 等ICLR 2026 · 被引用 87 次
相关 Paper
- Learning to Select Visual Tools from ExperienceZeyi Huang, Yuyang Ji, Anirudh Sundara Rajan, Zefan Cai 等CVPR 2026
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy OptimizationXinhai Hou, Shaoyuan Xu, Manan Biyani, Moyan Li 等CVPR 2026 · 被引用 25 次
- PyVision-RL: Forging Open Agentic Vision Models via RLShitian Zhao, Shaoheng Lin, Ming Li, Haoquan Zhang 等ICML 2026 · 被引用 10 次
- Current Agents Fail to Leverage World Model as Tool for ForesightCheng Qian, Emre Can Acikgoz, Bingxuan Li, Xiusi Chen 等ACL 2026 · 被引用 10 次
- Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced AgentsGuangfu Guo, Xiaoqian Lu, Yue FengEMNLP 2025
