VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
Yikun Wang, Siyin Wang, Qinyuan Cheng, Zhaoye Fei, Liang Ding, Qipeng Guo, Dacheng Tao, Xipeng Qiu
Abstract
Recent advancements in Large Vision-Language Models have showcased remarkable capabilities. However, they often falter when confronted with complex reasoning tasks that humans typically address through visual aids and deliberate, step-by-step thinking. While existing methods have explored text-based slow thinking or rudimentary visual assistance, they fall short of capturing the intricate, interleaved nature of human visual-verbal reasoning processes. To overcome these limitations and inspired by the mechanisms of slow thinking in human cognition, we introduce VisuoThink, a novel framework that seamlessly integrates visuospatial and linguistic domains. Visuo-Think facilitates multimodal slow thinking by enabling progressive visual-textual reasoning and incorporates test-time scaling through look-ahead tree search. Extensive experiments demonstrate that VisuoThink significantly enhances reasoning capabilities via inferencetime scaling, even without fine-tuning, achieving state-of-the-art performance in tasks involving geometry and spatial reasoning. Our code has been open-sourced at https: //github.com/ekonwang/VisuoThink .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b97a0c95-6b8a-40a4-a9a4-bdc39886b2b9Cited by top-tier papers6
- Look-Back: Implicit Visual Re-focusing in MLLM ReasoningShuo Yang, Yuwei Niu, Yuyang Liu, Yang Ye et al.AAAI 2026 · 32 citations
- Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary ConstructionsJingxuan Wei, Caijun Jia, Qi Chen, Honghao He et al.CVPR 2026 · 14 citations
- EcoAlign: An Economically Rational Framework for Efficient LVLM AlignmentRuoxi Cheng, Haoxuan Ma, Teng Ma, Hongyi ZhangCVPR 2026 · 6 citations
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGIHuanjin Yao, Jiaxing Huang, Yawen Qiu, Michael K. Chen et al.ICCV 2025 · 4 citations
- Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought ModelsJi Ma, Wei Suo, Peng Wang, Yanning ZhangCVPR 2026 · 3 citations
Builds on17
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
- Unleashing Text-to-Image Diffusion Models for Visual PerceptionWenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu et al.ICCV 2023 · 327 citations
- AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and TrainingZiyu Wan, Xidong Feng, Muning Wen, Stephen Marcus McAleer et al.ICML 2024 · 325 citations
- Self-Evaluation Guided Beam Search for ReasoningYuxi Xie, Kenji Kawaguchi, Yiran Zhao, James Xu Zhao et al.NeurIPS 2023 · 316 citations
Related papers
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited ViewsZhangquan Chen, Manyuan Zhang, Xinlei Yu, Xufang Luo et al.CVPR 2026 · 61 citations
- ProxyThinker: Test-Time Guidance through Small Visual ReasonersZilin Xiao, Jaywon Koo, Siru Ouyang, Jefferson Hernandez et al.ICLR 2026 · 8 citations
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL CyclesYihe Deng, Hritik Bansal, Fan Yin, Nanyun Peng et al.NeurIPS 2025 · 61 citations
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu et al.EMNLP 2025
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought ReasoningJiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li et al.ICLR 2026 · 51 citations
