Multi-LLM Debate: Framework, Principals, and Interventions
Andrew Estornell, Yang Liu
Abstract
The flexible and generalized nature of large language models has allowed for their application in a wide array of language-based domains. Much like their human contemporaries, these models are capable of engaging in discussions and debates as a means of improving answer quality. We first take a theoretical approach to analyzing debate and provide a framework through which debate can be mathe-matically examined. Building on this framework, we provide several theoretical results for multi-agent debate. In particular, we demonstrate that similar model capabilities, or similar model responses, can result in static debate dynamics where the debate procedure simply converges to the majority opinion. When this majority opinion is the result of a common misconception (possibly ingrained in the models through shared training data) debate is likely to converge to answers associated with that common misconception. Using insights from our theoretical results, we then propose three interventions that improve the efficacy of debate. For each intervention, we provide theoretical results demonstrating how debate is improved. We also demonstrate that these interventions result in better performance on four common benchmark tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7c2be81-8237-48be-b8db-56a4cb04d08aCited by top-tier papers13
- Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?Hyeong Kyu Choi, Xiaojin Zhu, Sharon LiNeurIPS 2025 · 93 citations
- Multi-Agent Debate for LLM Judges with Adaptive Stability DetectionTianyu Hu, Zhen Tan, Song Wang, Huaizhi Qu et al.NeurIPS 2025 · 25 citations
- Learning Decentralized LLM Collaboration with Multi-Agent Actor CriticShuo Liu, Tianle Chen, Ryan Amiri, Christopher AmatoICML 2026 · 6 citations
- The Value of Variance: Mitigating Debate Collapse in Multi-Agent Systems via Uncertainty-Driven Policy OptimizationLuoxi Tang, Yuqiao Meng, Joseph Costa, Yingxue Zhang et al.ICML 2026 · 4 citations
- Self Iterative Label Refinement via Robust Unlabeled LearningHikaru Asano, Tadashi Kozuno, Yukino BabaNeurIPS 2025 · 1 citation
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
Related papers
- Multi-Agent Debate with Memory MaskingHongduan Tian, Xiao Feng, Ziyuan Zhao, Xiangyu Zhu et al.ICLR 2026 · 7 citations
- Beyond Correctness: Distance-Based Social Dynamics of Multi-Agent DebateSeungwoong Ha, Melanie MitchellICML 2026
- Breaking Mental Set to Improve Reasoning through Diverse Multi-Agent DebateYexiang Liu, Jie Cao, Zekun Li, Ran He et al.ICLR 2025
- Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMsAndries P. Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D. Barrett et al.ICML 2024 · 82 citations
- Analyze-Compose-Execute: A Dynamic Dialogue Framework for Multi-Agent DebateWenyuan Gu, Haowen Wang, Jiale Han, Xiang Li et al.AAAI 2026
