Why Keep Your Doubts to Yourself? Trading Visual Uncertainties among Vision-Language Models
Jusheng Zhang, Yijia Fan, Kaitong Cai, Jing Yang, Jiawei Yao, Jian Wang, Guanlong Qu, Ziliang Chen, Keze Wang
Abstract
Vision-Language Models (VLMs) enable powerful multi-agent systems, but scaling them is economically unsustainable: coordinating heterogeneous agents under information asymmetry often spirals costs. Existing paradigms, such as Mixture-of-Agents and knowledge-based routers, rely on heuristic proxies that ignore costs and collapse uncertainty structure, leading to provably suboptimal coordination. We introduce Agora, a framework that reframes coordination as a decentralized market for uncertainty. Agora formalizes epistemic uncertainty into a structured, tradable asset (perceptual, semantic, inferential), and enforces profitability-driven trading among agents based on rational economic rules. A market-aware broker, extending Thompson Sampling, initiates collaboration and guides the system toward cost-efficient equilibria. Experiments on five multimodal benchmarks (MMMU, MMBench, MathVision, InfoVQA, CC-OCR) show that Agora outperforms strong VLMs and heuristic multi-agent strategies, e.g., achieving +8.5% accuracy over the best baseline on MMMU while reducing cost by over 3×. These results establish market-based coordination as a principled and scalable paradigm for building economically viable multi-agent visual intelligence systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c87f92ec-f51a-4f95-8260-be53a33b59f6Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim et al.ICLR 2024 · 282 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
Related papers
- ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question AnsweringAymen Lassoued, Mohamed Ali Souibgui, Yousri KessentiniCVPR 2026 · 1 citation
- Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time ScalingXinlei Yu, Chengming Xu, Zhangquan Chen, Yudong Zhang et al.CVPR 2026 · 31 citations
- GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual ReasoningJusheng Zhang, Yijia Fan, Wenjun Lin, Ruiqi Chen et al.NeurIPS 2025 · 75 citations
- KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent SystemsJusheng Zhang, Zimeng Huang, Yijia Fan, Ningyuan Liu et al.ICML 2025
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent EnvironmentsZelai Xu, Zhexuan Xu, Xiangmin Yi, Huining Yuan et al.CVPR 2026 · 3 citations
