M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, Wanxiang Che
摘要
Multi-modal Chain-of-Thought (MCoT) requires models to leverage knowledge from both textual and visual modalities for step-bystep reasoning, which gains increasing attention. Nevertheless, the current MCoT benchmark still faces some challenges: (1) absence of visual modal reasoning, (2) single-step visual modal reasoning, and (3) Domain missing, thereby hindering the development of MCoT. Motivated by this, we introduce a novel benchmark (M 3 CoT) to address the above challenges, advancing the multi-domain, multi-step, and multi-modal CoT. Additionally, we conduct a thorough evaluation involving abundant MCoT approaches on Vision Large Language Models (VLLMs). In addition, we highlight that the current VLLMs still struggle to correctly reason in M 3 CoT and there remains a large gap between existing VLLMs and human performance in M 3 CoT, despite their superior results on previous MCoT benchmarks. To our knowledge, we take the first meaningful step toward the multi-domain, multi-step, and multi-modal scenario in MCoT. We hope that M 3 CoT can serve as a valuable resource, providing a pioneering foundation in multi-domain, multi-step, multi-modal chain-of-thought research. * Corresponding Author Q : … supports the plant … Which part do we usually eat? A: (B) the stem O: … (B) Only to indicate the time A: (B) soft A: (C) To indicate … R: … The feather is soft… (b) Single-step visual modal reasoning. (c) Multi-step visual modal reasoning. Q: Which property matches this object? O: … (B) soft O: …(B) stem R: Step 1: The wind vane on top … indicate the wind direction. Step 2: …. The clock on top …it is used to indicate the time. VLLM VLLM VLLM (a) Absence of visual modal reasoning. R: … we usually eat is the stem. It supports the plant … Single Step Missing Multi-Step 1 Multi-Step 2 Q : What is the purpose of the tower? (C) To indicate the time and wind direction…
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-ThoughtQiguang Chen, Libo Qin, Jiaqi Wang, Jingxuan Zhou 等NeurIPS 2024 · 被引用 104 次
- Visual Planning: Let's Think Only with ImagesYi Xu, Chengzu Li, Han Zhou, Xingchen Wan 等ICLR 2026 · 被引用 93 次
- Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-ThoughtZihui Cheng, Qiguang Chen, Xiao Xu, Jiaqi Wang 等NeurIPS 2025 · 被引用 38 次
- What Factors Affect Multi-Modal In-Context Learning? An In-Depth ExplorationLibo Qin, Qiguang Chen, Hao Fei, Zhi Chen 等NeurIPS 2024 · 被引用 37 次
- From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal ReasoningRuilin Luo, Chufan Shi, Yizhen Zhang, Cheng Yang 等ICLR 2026 · 被引用 10 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language ModelsZihui Cheng, Qiguang Chen, Jin Zhang, Hao Fei 等AAAI 2025 · 被引用 36 次
- M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image UnderstandingJuntao Jiang, Jiangning Zhang, Yali Bi, Jinsheng Bai 等ICLR 2026 · 被引用 3 次
- M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question AnsweringJiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao 等ACL 2026
- Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot DoZhuoran Jin, Kejian Zhu, Hongbang Yuan, Yupu Hao 等ACL 2026
- Interleaved-Modal Chain-of-ThoughtJun Gao, Yongqi Li, Ziqiang Cao, Wenjie LiCVPR 2025
