CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
Guanghao Zhang, Tao Zhong, Yan Xia, Mushui Liu, Zhelun Yu, Haoyuan Li, Wanggui He, Dong She, Yi Wang, Hao Jiang
Abstract
While previous multimodal slow-thinking methods have demonstrated remarkable success in single-image understanding scenarios, their effectiveness becomes fundamentally constrained when extended to more complex multiimage comprehension tasks. This limitation stems from their predominant reliance on text-based intermediate reasoning processes. While humans, when engaging in sophisticated multi-image analysis, typically perform two complementary cognitive operations: (1) continuous cross-image visual comparison through region-of-interest matching, and (2) dynamic memorization of critical visual concepts throughout the reasoning chain. Motivated by these observations, we propose the Complex Multi-Modal Chain-of-Thought (CMMCoT) framework, a multi-step reasoning framework that mimics human-like "slow thinking" for multi-image understanding. Our approach incorporates two key innovations: (1) The construction of interleaved multimodal multi-step reasoning chains, which utilize critical visual region tokens, extracted from intermediate reasoning steps, as supervisory signals. This mechanism not only facilitates comprehensive crossmodal understanding but also enhances model interpretability. (2) The introduction of a test-time memory augmentation module that expands the model's reasoning capacity during inference while preserving parameter efficiency. Furthermore, to facilitate research in this direction, we have curated a novel multi-image slow-thinking dataset. Extensive experiments demonstrate the effectiveness of our model. Code is available at https://github.com/zhangguanghao523/ CMMCoT. Recent years have witnessed the rapid advancement of generative (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Latent Visual ReasoningBangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang et al.ICLR 2026 · 80 citations
- Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language ModelsJiaqi Liu, Lang Sun, Ronghao Fu, Bo YangICLR 2026 · 22 citations
- Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary ConstructionsJingxuan Wei, Caijun Jia, Qi Chen, Honghao He et al.CVPR 2026 · 14 citations
- Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMsXudong Li, Mengdan Zhang, Peixian Chen, Xiawu Zheng et al.NeurIPS 2025 · 4 citations
- OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language ModelsQiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu et al.ACL 2026 · 1 citation
Builds on16
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
- Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level VisionHaoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen et al.ICLR 2024 · 258 citations
Related papers
- VGR: Visual Grounded ReasoningJiacong Wang, Zijian Kang, Haochen Wang, Xiao Liang et al.ICLR 2026 · 64 citations
- Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and VisionLuozheng Qin, Jia Gong, Yuqing Sun, Tianjiao Li et al.ICLR 2026 · 55 citations
- DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language ModelsGe Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou et al.NeurIPS 2023 · 252 citations
- VisuoThink: Empowering LVLM Reasoning with Multimodal Tree SearchYikun Wang, Siyin Wang, Qinyuan Cheng, Zhaoye Fei et al.ACL 2025 · 35 citations
- Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot DoZhuoran Jin, Kejian Zhu, Hongbang Yuan, Yupu Hao et al.ACL 2026
