VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
Mingyuan Wu, Jingcheng Yang, Jize Jiang, Meitang Li, Kaizhuo Yan, Hanchao Yu, Minjia Zhang, ChengXiang Zhai, Klara Nahrstedt
Abstract
Reinforcement learning finetuning (RFT) has significantly advanced the reasoning capabilities of large language models (LLMs) by enabling long chains of thought, multi-turn self-correction, and effective tool use. While recent works attempt to extend RFT to vision-language models (VLMs), these efforts largely focus on textonly reasoning conditioned on original image inputs, and do not incorporate visual reasoning in the response. In contrast, test-time methods like Visual Sketchpad incorporate visual steps but lack training mechanisms. We introduce VTool-R1, the first RFT framework that trains VLMs to generate multimodal chains of thought by interleaving text and intermediate visual reasoning steps. VTool-R1 integrates Python-based visual editing tools into the RFT process, enabling VLMs to learn when and how to generate visual reasoning steps that enhance the final output quality. Trained with outcome-based rewards, our approach elicits strategic visual tool use for multi-modal reasoning without relying on process-based supervision. Extensive experiments on structured visual reasoning over charts and tables show that VTool-R1 enhances reasoning performance by teaching VLMs to "think with images" and generate multimodal chain of thoughts with tools. To support future research in multi-turn multi-modal reasoning, we open-source our code at https://github.com/VTOOL-R1/vtool-r1 . * Mingyuan and Jingcheng contributed equally.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94c24cc8-9812-4cb9-a184-dd465a587ed2Cited by top-tier papers18
- MemGen: Weaving Generative Latent Memory for Self-Evolving AgentsGuibin Zhang, Muxin Fu, Shuicheng YanICLR 2026 · 102 citations
- Latent Visual ReasoningBangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang et al.ICLR 2026 · 80 citations
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited ViewsZhangquan Chen, Manyuan Zhang, Xinlei Yu, Xufang Luo et al.CVPR 2026 · 61 citations
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RLSiyi Chen, Mikaela Angelina Uy, Chan Hee Song, Faisal Ladhak et al.CVPR 2026 · 24 citations
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMsMeng Lu, Ran Xu, Yi Fang, Wenxuan Zhang et al.CVPR 2026 · 15 citations
Builds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
Related papers
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL CyclesYihe Deng, Hritik Bansal, Fan Yin, Nanyun Peng et al.NeurIPS 2025 · 61 citations
- Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement FinetuningMinheng Ni, Zhengyuan Yang, Linjie Li, Chung-Ching Lin et al.NeurIPS 2025 · 35 citations
- Monet: Reasoning in Latent Visual Space Beyond Image and LanguageQixun Wang, Yang Shi, Yifei Wang, Yuanxing Zhang et al.CVPR 2026
- ReFocus: Visual Editing as a Chain of Thought for Structured Image UnderstandingXingyu Fu, Minqian Liu, Zhengyuan Yang, John Corring et al.ICML 2025
