CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
Xinhai Hou, Shaoyuan Xu, Manan Biyani, Moyan Li, Jia Liu, Todd C. Hollon, Bryan Wang
Abstract
Agentic vision-language models are increasingly trained to "think with images" by calling image operations. However, we show that high final-answer accuracy often hides unfaithful visual reasoning: models may invoke tools on irrelevant regions or ignore tool outputs entirely, yet still guess the correct answer. In this work, we first propose a faithfulness evaluation protocol that measures whether intermediate visual tool outputs (e.g., crops) actually contain the queried evidence. This reveals that recent visual agents achieve high final-answer accuracy but exhibit low rates of faithful tool-use on visual search benchmarks. We then introduce CodeV, a code-based visual agent trained with Tool-Aware Policy Optimization (TAPO). TAPO is a processlevel RL framework that augments GRPO with dense rewards defined directly on visual tool inputs and outputs, rather than on chain-of-thought tokens, making supervision easier to verify and less susceptible to reward hacking. CodeV represents visual tools as executable Python code, and TAPO assigns step-wise rewards based solely on the question and tool output, encouraging both necessary and evidence-consistent tool use. In a two-stage SFT+RL pipeline, CodeV achieves competitive or superior accuracy while substantially increasing faithful tool-use rates on related visual search benchmarks. Beyond visual search, CodeV attains strong performance on a range of multimodal reasoning and math benchmarks, suggesting that explicitly supervising intermediate tool behavior is crucial for building trustworthy, agentic visual reasoning systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd62af4a-2cd8-47a3-8787-5d7777e94c63Cited by top-tier papers4
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal PerceptionLai Wei, Liangbo He, jun lan, Lingzhong Dong et al.ICML 2026 · 27 citations
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric ReasoningXiang Fang, Wanlong Fang, Changshuo WangCVPR 2026 · 17 citations
- PyVision-RL: Forging Open Agentic Vision Models via RLShitian Zhao, Shaoheng Lin, Ming Li, Haoquan Zhang et al.ICML 2026 · 10 citations
- Learning Transferable Temporal Primitives for Video Reasoning via Synthetic VideosSongtao Jiang, Sibo Song, Chenyi Zhou, Yuan Wang et al.CVPR 2026 · 3 citations
Builds on18
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang et al.ICLR 2026 · 406 citations
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
Related papers
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual ReasoningQi Song, Honglin Li, Yingchen Yu, Haoyi Zhou et al.CVPR 2026 · 16 citations
- Empowering LLM Tool Invocation with Tool-call Reward ModelDa Ma, Ziyue Yang, Hongshen Xu, Haotian Fang et al.ICLR 2026
- CFPO: Counterfactual Policy Optimization for Multimodal ReasoningZhangyuan Yu, Wanran Sun, Guangjing Yang, Xiaohu Wu et al.ICML 2026
- Thinking with Programming Vision: Towards a Unified View for Thinking with ImagesZirun Guo, Minjie Hong, Feng Zhang, Kai Jia et al.CVPR 2026 · 15 citations
- Visually-Guided Policy Optimization for Multimodal ReasoningZengbin Wang, Feng Xiong, Liang Lin, Xuecai Hu et al.ACL 2026 · 7 citations
