A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning
Changmeng Zheng, Dayong Liang, Wengyu Zhang, Xiaoyong Wei, Tat-Seng Chua, Qing Li
Abstract
This paper presents a pilot study aimed at introducing multi-agent debate into multimodal reasoning. The study addresses two key challenges: the trivialization of opinions resulting from excessive summarization and the diversion of focus caused by distractor concepts introduced from images. These challenges stem from the inductive (bottom-up) nature of existing debating schemes. To address the issue, we propose a deductive (top-down) debating approach called Blueprint Debate on Graphs (BDoG). In BDoG, debates are confined to a blueprint graph to prevent opinion trivialization through world-level summarization. Moreover, by storing evidence in branches within the graph, BDoG mitigates distractions caused by frequent but irrelevant concepts. Extensive experiments validate that BDoG is able to achieve state-of-the-art results in ScienceQA and MMBench with significant improvements over previous methods. The source code can be accessed at https://github.com/thecharm/BDoG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0a60f4c-7209-4cab-90ff-cfca8a6cf6f6Cited by top-tier papers9
- MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLMBowen Dong, Minheng Ni, Zitong Huang, Guanglei Yang et al.NeurIPS 2025 · 25 citations
- Few-Shot Joint Multimodal Entity-Relation Extraction via Knowledge-Enhanced Cross-modal Prompt ModelLi Yuan, Yi Cai, Junsheng HuangACM MM 2024 · 9 citations
- Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-SpoofingHonglu Zhang, Zhiqin Fang, Ningning Zhao, Saihui Hou et al.CVPR 2026 · 4 citations
- Multi-Agent Undercover Gaming: Hallucination Removal Through Counterfactual Test for Multimodal ReasoningDayong Liang, Xiao-Yong Wei, Changmeng ZhengAAAI 2026 · 1 citation
- Beyond Single-View Detection: A Dual-Space Reasoning Framework for Interpretable Harmful Meme UnderstandingWenqing Hou, Hongkui Tu, Ye Wang, Yue Zhang et al.ACL 2026
Builds on26
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
Related papers
- MAD-Logic: Multi-Agent Debate Enhances Symbolic Translation and ReasoningHaocheng Yang, Fengxiang Cheng, Tianjun Yao, Mengyue Yang et al.ICLR 2026
- iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM InferenceWei Fan, JinYi Yoon, Bo JiAAAI 2026 · 5 citations
- Reasoning on Knowledge Graphs with Debate DynamicsMarcel Hildebrandt, Jorge Andres Quintero Serna, Yunpu Ma, Martin Ringsquandl et al.AAAI 2020 · 59 citations
- MAR: Metacognitive Agentic Reasoning for Multimodal Fake News DetectionWenyu Chen, Hengbing Dong, Junhao Wa, Ping Wei et al.KDD 2026
- Debate on Graph: A Flexible and Reliable Reasoning Framework for Large Language ModelsJie Ma, Zhitao Gao, Qi Chai, Wangchun Sun et al.AAAI 2025 · 8 citations
