Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
Timin Gao, Peixian Chen, Mengdan Zhang, Chaoyou Fu, Yunhang Shen, Yan Zhang, Shengchuan Zhang, Xiawu Zheng, Xing Sun, Liujuan Cao, Rongrong Ji
Abstract
With the advent of large language models(LLMs) enhanced by the chain-of-thought(CoT) methodology, the visual reasoning problem is usually decomposed into manageable sub-tasks and tackled sequentially with various external tools. However, such a paradigm faces the challenge of the potential "determining hallucinations" in decision generation due to insufficient visual information and the limitation of low-level perception tools that fail to provide abstract summaries necessary for comprehensive reasoning. We argue that converging visual context acquisition and logical reasoning is pivotal for tackling visual reasoning tasks. This paper delves into the realm of multimodal CoT to solve intricate visual reasoning tasks with multimodal large language models(MLLMs) and their cognitive capability. To this end, we propose an innovative multimodal CoT framework, termed Cantor, characterized by a perception-decision architecture. Cantor first acts as a decision generator and integrates visual inputs to analyze the image and problem, ensuring a closer alignment with the actual context. Furthermore, Cantor leverages the advanced cognitive functions of MLLMs to perform as multifaceted experts for deriving higher-level information, enhancing the CoT generation process. Our extensive experiments demonstrate the efficacy of the proposed framework, showing significant improvements in multimodal CoT performance across two complex visual reasoning datasets, without necessitating fine-tuning or ground-truth rationales. Project Page: https://ggg0919.github.io/cantor/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2c0f9c7-2057-4497-9b29-1b5a47f84deaCited by top-tier papers31
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual TokensZeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen et al.CVPR 2026 · 124 citations
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial ReasoningYang Liu, Ming Ma, Xiaomin Yu, Pengxiang Ding et al.NeurIPS 2025 · 46 citations
- VisuoThink: Empowering LVLM Reasoning with Multimodal Tree SearchYikun Wang, Siyin Wang, Qinyuan Cheng, Zhaoye Fei et al.ACL 2025 · 35 citations
- LOVA3: Learning to Visual Question Answering, Asking and AssessmentHenry Hengyuan Zhao, Pan Zhou, Difei Gao, Zechen Bai et al.NeurIPS 2024 · 30 citations
- Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score CollaborationZhixuan Shen, Haonan Luo, Kexun Chen, Fengmao Lv et al.AAAI 2025 · 21 citations
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Chain-of-Thought Guided Multi-Modal Object Re-IdentificationYa Gao, Shihao Li, Zhaojun Liu, Aihua Zheng et al.CVPR 2026
- Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot DoZhuoran Jin, Kejian Zhu, Hongbang Yuan, Yupu Hao et al.ACL 2026
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu et al.EMNLP 2025
- Unsupervised Visual Chain-of-Thought Reasoning via Preference OptimizationKesen Zhao, Beier Zhu, Qianru Sun, Hanwang ZhangICCV 2025 · 3 citations
- See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMsYongchang Zhang, Xianzheng Ma, Tianyi Liu, Guangquan Zhou et al.CVPR 2026 · 2 citations
