Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
Zihui Cheng, Qiguang Chen, Xiao Xu, Jiaqi Wang, Weiyun Wang, Hao Fei, Yidong Wang, Alex Jinpeng Wang, Zhi Chen, Wanxiang Che, Libo Qin
Abstract
Large Vision-Language Models (LVLMs) have achieved significant success in multimodal tasks, with multimodal chain-of-thought (MCoT) further enhancing performance and interpretability. Recent MCoT methods fall into two categories: (i) Textual-MCoT (T-MCoT), which takes multimodal input and produces textual output; and (ii) Interleaved-MCoT (I-MCoT), which generates interleaved image-text outputs. Despite advances in both approaches, the mechanisms driving these improvements are not fully understood. To fill this gap, we first reveal that MCoT boosts LVLMs by incorporating visual thoughts, which convey image information to the reasoning process regardless of the MCoT format, depending only on clarity and conciseness of expression. Furthermore, to explore visual thoughts systematically, we define four distinct forms of visual thought expressions and analyze them comprehensively. Our findings demonstrate that these forms differ in clarity and conciseness, yielding varying levels of MCoT improvement. Additionally, we explore the internal nature of visual thoughts, finding that visual thoughts serve as intermediaries between the input image and reasoning to deeper transformer layers, enabling more advanced visual information transmission. We hope that the visual thoughts can inspire further breakthroughs for future MCoT research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0a44cb0-dfba-4d20-bf35-5baa48f1fd41Cited by top-tier papers16
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual TokensZeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen et al.CVPR 2026 · 124 citations
- InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual SearchKaican Li, Lewei Yao, Jiannan Wu, Tiezheng YU et al.ICLR 2026 · 10 citations
- DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid ThinkingWeicheng Zheng, Xiaofei Mao, Nanfei Ye, Pengxiang Li et al.ICLR 2026 · 8 citations
- FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-and-Language NavigationJing Zuo, Lingzhou Mu, Fan Jiang, Chengcheng Ma et al.CVPR 2026 · 4 citations
- ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language ModelsYongheng Zhang, Xu Liu, Ruihan Tao, Qiguang Chen et al.ACM MM 2025 · 4 citations
Builds on28
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language ModelsZihui Cheng, Qiguang Chen, Jin Zhang, Hao Fei et al.AAAI 2025 · 36 citations
- Interleaved-Modal Chain-of-ThoughtJun Gao, Yongqi Li, Ziqiang Cao, Wenjie LiCVPR 2025
- Imagine While Reasoning in Space: Multimodal Visualization-of-ThoughtChengzu Li, Wenshan Wu, Huanyu Zhang, Yan Xia et al.ICML 2025
- Rationale-Enhanced Decoding for Multi-modal Chain-of-ThoughtShin'ya Yamaguchi, Kosuke Nishida, Daiki ChijiwaCVPR 2026
- Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought ModelsJi Ma, Wei Suo, Peng Wang, Yanning ZhangCVPR 2026 · 3 citations
