Grounded Chain-of-Thought for Multimodal Large Language Models
Qiong Wu, Xiangcong Yang, Yiyi Zhou, Chenxin Fang, Baiyang Song, Xiaoshuai Sun, Rongrong Ji
摘要
Despite great progress, existing multimodal large language models (MLLMs) are prone to visual hallucination, greatly impeding their trustworthy applications. In this paper, we study this problem from the perspective of visual-spatial reasoning, and propose a new learning task for MLLMs, termed Grounded Chain-of-Thought (GCoT). Different from recent visual CoT studies, which focus more on visual knowledge reasoning, GCoT is keen to helping MLLMs to recognize and ground the relevant visual cues step by step, thereby predicting the correct answer with grounding coordinates as the intuitive basis. To facilitate this task, we also carefully design and construct a dataset called multimodal grounded chain-of-thought (MM-GCoT) consisting of 24,022 GCoT examples for 5,033 images. Besides, a comprehensive consistency evaluation system is also introduced, including the metrics of answer accuracy , grounding accuracy and answer-grounding consistency. We further design and conduct a bunch of experiments on 12 advanced MLLMs, and reveal some notable findings: i. most MLLMs performs poorly on the consistency evaluation, indicating obvious visual hallucination; ii., visual hallucination is not directly related to the parameter size and general multimodal performance, i.e., a larger and stronger MLLM is not less affected by this issue. Lastly, we also demonstrate that the proposed dataset can help existing MLLMs to well cultivate their GCoT capability and reduce the inconsistent answering significantly. Moreover, their GCoT can be also generalized to exiting multimodal tasks, such as open-world QA and REC. Our dataset and evaluation scripts are anonymously released at: https: //github.com/DoubtedSteam/MM-GCoT
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language ReasoningJiaqi Liu, Kaiwen Xiong, Peng Xia, Yiyang Zhou 等ICML 2026 · 被引用 28 次
- Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic ShuffleLinghao Zhu, Yiran Guan, Dingkang Liang, Jianzhong Ju 等ICLR 2026 · 被引用 18 次
- SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual ScenesChuhan Wang, Xintong Li, Jennifer Yuntong Zhang, Junda Wu 等ACL 2026 · 被引用 9 次
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual ReasoningZejun Li, Yingxiu Zhao, Jiwen Zhang, Siyuan Wang 等ICLR 2026 · 被引用 7 次
- Can Vision-Language Models Assess Graphic Design Aesthetics? A Benchmark, Evaluation, and Dataset PerspectiveRuichuan An, Shizhao Sun, Danqing Huang, Mingxi Cheng 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper21
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought ModelsJi Ma, Wei Suo, Peng Wang, Yanning ZhangCVPR 2026 · 被引用 3 次
- See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMsYongchang Zhang, Xianzheng Ma, Tianyi Liu, Guangquan Zhou 等CVPR 2026 · 被引用 2 次
- Mitigating Low-Quality Reasoning in MLLMs: Self-Driven Refined Multimodal CoT with Selective Thinking and Step-wise Visual EnhancementChongjun Tu, Peng Ye, Dongzhan Zhou, Tao Chen 等AAAI 2026
- An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal ModelsFatemeh Shiri, Xiao-Yu Guo, Mona Far, Xin Yu 等EMNLP 2024 · 被引用 7 次
- Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot DoZhuoran Jin, Kejian Zhu, Hongbang Yuan, Yupu Hao 等ACL 2026
