GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning
Jusheng Zhang, Yijia Fan, Wenjun Lin, Ruiqi Chen, Haoyi Jiang, Wenhao Chai, Jian Wang, Keze Wang
Abstract
We propose GAM-Agent, a game-theoretic multi-agent framework for enhancing vision-language reasoning. Unlike prior single-agent or monolithic models, GAM-Agent formulates the reasoning process as a non-zero-sum game between base agents-each specializing in visual perception subtasks-and a critical agent that verifies logic consistency and factual correctness. Agents communicate via structured claims, evidence, and uncertainty estimates. The framework introduces an uncertainty-aware controller to dynamically adjust agent collaboration, triggering multi-round debates when disagreement or ambiguity is detected. This process yields more robust and interpretable predictions. Experiments on four challenging benchmarks-MMMU, MMBench, MVBench, and V*Bench-demonstrate that GAM-Agent significantly improves performance across various VLM backbones. Notably, GAM-Agent boosts the accuracy of small-to-mid scale models (e.g., Qwen2.5-VL-7B, InternVL3-14B) by 5-6%, and still enhances strong models like GPT-4o by up to 2-3%. Our approach is modular, scalable, and generalizable, offering a path toward reliable and explainable multi-agent multimodal reasoning.
However, current collaborative game-theoretic methods are often highly complex and heavily rely on reasoning paths, clues, and the fusion of multiple information, making them difficult to apply directly 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
to visual reasoning [54,14,12,72,48]. To address this issue, this paper explores the underlying architecture of existing VLMs, effectively utilizes important intermediate results in the reasoning process, and extracts representations of uncertainty in reasoning outcomes [21]. Specifically, we propose a collaborative game-theoretic framework [37,26,44,30], named GAM-Agent, based on game-theoretic and uncertainty-aware inference. In this way, the complex visual reasoning process can be modeled as a non-zero-sum game involving multiple agents collaborating to reach a consensus [63,26]. Specifically, we encourage agents to share their respective assessments of reasoning uncertainty and engage in a progressive interactive game to guide GAM-Agent to evolve step-by-step and ultimately reach a consensus.
To address these challenges, we introduce GAM-Agent, a novel agent collaboration framework centered around a strategic interplay between two specialized agent cohorts: Base Agents and Critical Agents. The Base Agents are tasked with initial visual interpretation and evidence generation from distinct perspectives, such as object recognition, scene description, and textual analysis from images. Concurrently, Critical Agents, acting as reasoning critique experts, scrutinize the outputs from Base Agents and evaluate factual accuracy, logical coherence, and overall completeness. The core of our GAM-Agent lies in modeling the interaction between these Base and Critical Agents as a nonzero-sum game, fundamentally arbitrated by quantified uncertainty. In this game, agents iteratively share and refine their uncertainty assessments regarding their claims and evidence, engaging in a structured debate process aimed at progressively reducing ambiguity and converging toward a consensus. This uncertainty-driven, game-theoretic collaboration allows for dynamic and strategic integration of diverse insights, leading to more robust and reliable visual reasoning outcomes. Specifically, Base Agents first generate diverse preliminary analyses and identify supporting evidence for their claims. These outputs are then processed by a Claim Parser module, which deconstructs the unstructured responses into structured information tuples, and an Evidence Mapping module, which links these textual claims to specific visual regions in the input image, thereby grounding the reasoning process. Besides, an Uncertainty Quantification mechanism continuously assesses the confidence of each agent's contribution. The entire process is orchestrated by a Debate Controller & Integrator. This component first evaluates the initial consensus and system uncertainty. If significant discrepancies or high uncertainty are detected, it initiates an iterative debate. During this debate, Base Agents refine their arguments, while Critical Agents provide evaluations, with the Uncertainty Quantification guiding the dynamic weighting and integration of information. This iterative loop continues, progressively refining the collective understanding and reducing uncertainty until a robust consensus is achieved or termination criteria are met. Extensive and comprehensive evaluations on large-scale benchmarks demonstrate the superiority of our GAM-Agent. Experimental results show that GAM-Agent achieves significant performance improvements on multiple complex visual reasoning benchmarks such as MMMU, MMBench, MVBench, and V*Bench. For example, it boosts the accuracy of small-to-mid scale VLMs by 5-6% and enhances the accuracy of top-tier models like GPT-4o by up to 2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & SparsificationWei Li, Renshan Zhang, Rui Shao, Jie He et al.NeurIPS 2025 · 87 citations
- CF-VLM: CounterFactual Vision-Language Fine-tuningJusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang et al.NeurIPS 2025 · 71 citations
- MAT-Agent: Adaptive Multi-Agent Training OptimizationJusheng Zhang, Kaitong Cai, Yijia Fan, Ningyuan Liu et al.NeurIPS 2025 · 46 citations
- FastVMT: Eliminating Redundancy in Video Motion TransferYue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng et al.ICLR 2026 · 32 citations
- Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region ControlZeqian Long, Mingzhe Zheng, Kunyu Feng, Xinhua Zhang et al.ICLR 2026 · 29 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
Related papers
- MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical ReasoningPeng Xia, Jinglu Wang, Yibo Peng, Kaide Zeng et al.ICLR 2026 · 47 citations
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai et al.AAAI 2026
- Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time ScalingXinlei Yu, Chengming Xu, Zhangquan Chen, Yudong Zhang et al.CVPR 2026 · 31 citations
- LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM AgentsBoyu Chen, Zhengrong Yue, Siran Chen, Zikang Wang et al.ICCV 2025 · 12 citations
- ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question AnsweringAymen Lassoued, Mohamed Ali Souibgui, Yousri KessentiniCVPR 2026 · 1 citation
