Don’t Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs
Muhammad Kamran Janjua, Hugo Silva, Di Niu, Bahador Rashidi
Abstract
Multimodal language models (MLLMs) are increasingly paired with vision tools (e.g., depth, flow, correspondence) to enhance visual reasoning. However, despite access to these tool-generated visual cues, MLLMs often fail to benefit from them. Existing approaches typically feed raw tool outputs into the model, but these dense, pixel-level representations are misaligned with the language-native reasoning strengths of LLMs, leading to weak perception and reliance on language priors. We argue that, in problems where vision tools can provide the necessary visual cues, the bottleneck is not more tool calls or larger MLLMs, it is how tool outputs are represented. We introduce Perception Programs (P^2), a training-free, model-agnostic method that rewrites tool outputs into compact, structured, language-native summaries that MLLMs can directly parse and reason over. Across six perception-centric tasks in BLINK, P^2 consistently yields large improvements over base models and raw tool-augmented baselines. With GPT-5 Mini as the base model, P^2 raises its accuracy from 41.35% to 86.47% on multi-view reasoning, from 52.42% to 81.45% on relative depth, and achieves a 22% average gain across tasks, setting new state-of-the-art results. Even on smaller MLLMs, e.g., InternVL3.5-4B and Qwen3VL-4B, we observe 15–40% absolute gains from P^2, surpassing prior agentic, supervised, and RL-based tool-use methods—without any training or model modifications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76cedf08-4dd7-4148-80a6-a89af3842b87Builds on16
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with LanguageAndy Zeng, Maria Attarian, Brian Ichter, Krzysztof Marcin Choromanski et al.ICLR 2023 · 171 citations
- Thyme: Think Beyond ImagesYifan Zhang, Xingyu Lu, Shukang Yin, Chaoyou Fu et al.ICLR 2026 · 146 citations
Related papers
- Perception Tokens Enhance Visual Reasoning in Multimodal Language ModelsMahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh, Ethan Shen et al.CVPR 2025
- Reasoning-Aligned Perception Decoupling for Scalable Multi-modal ReasoningYunhao Gou, Kai Chen, Zhili Liu, Lanqing HONG et al.ICLR 2026 · 7 citations
- pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial ReasoningZhanpeng Luo, Ce Zhang, Silong Yong, Cunxi Dai et al.ICLR 2026 · 15 citations
- Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement TrainingQihuang Zhong, Liang Ding, Wenjie Xuan, Juhua Liu et al.ICML 2026 · 1 citation
- Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMsFangrui Zhu, Hanhui Wang, Yiming Xie, Jing Gu et al.NeurIPS 2025 · 7 citations
