Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
Andries P. Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D. Barrett, Arnu Pretorius
Abstract
Recent advancements in large language models (LLMs) underscore their potential for responding to inquiries in various domains. However, ensuring that generative agents provide accurate and reliable answers remains an ongoing challenge. In this context, multi-agent debate (MAD) has emerged as a promising strategy for enhancing the truthfulness of LLMs. We benchmark a range of debating and prompting strategies to explore the trade-offs between cost, time, and accuracy. Importantly, we find that multi-agent debating systems, in their current form, do not reliably outperform other proposed prompting strategies, such as self-consistency and ensembling using multiple reasoning paths. However, when performing hyperparameter tuning, several MAD systems, such as Multi-Persona, perform better. This suggests that MAD protocols might not be inherently worse than other approaches, but that they are more sensitive to different hyperparameter settings and difficult to optimize. We build on these results to offer insights into improving debating strategies, such as adjusting agent agreement levels, which can significantly enhance performance and even surpass all other non-debate protocols we evaluated. We provide an open-source repository to the community with several state-of-the-art protocols together with evaluation scripts to benchmark across popular research datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b06ae6e6-4a16-43e4-abff-66092a8de59dCited by top-tier papers24
- Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?Hyeong Kyu Choi, Xiaojin Zhu, Sharon LiNeurIPS 2025 · 93 citations
- SiriuS: Self-improving Multi-agent Systems via Bootstrapped ReasoningWanjia Zhao, Mert Yüksekgönül, Shirley Wu, James Y. ZouNeurIPS 2025 · 49 citations
- Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM SystemsShangbin Feng, Zifeng Wang, Palash Goyal, Yike Wang et al.NeurIPS 2025 · 26 citations
- Do We Truly Need So Many Samples? Multi-LLM Repeated Sampling Efficiently Scales Test-Time ComputeJianhao Chen, Zishuo Xun, Bocheng Zhou, Han Qi et al.AAAI 2026 · 18 citations
- Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability TheoryYexiang Liu, Zekun Li, Zhi Fang, Nan Xu et al.ACL 2025 · 12 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- Breaking Mental Set to Improve Reasoning through Diverse Multi-Agent DebateYexiang Liu, Jie Cao, Zekun Li, Ran He et al.ICLR 2025
- iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM InferenceWei Fan, JinYi Yoon, Bo JiAAAI 2026 · 5 citations
- Key Decision-Makers in Multi-Agent Debates: Who Holds the Power?Qian Zhang, Jinyi Liu, Yan Zheng, Hebin Liang et al.AAAI 2026
- MAD-Logic: Multi-Agent Debate Enhances Symbolic Translation and ReasoningHaocheng Yang, Fengxiang Cheng, Tianjun Yao, Mengyue Yang et al.ICLR 2026
- Multi-Agent Debate with Memory MaskingHongduan Tian, Xiao Feng, Ziyuan Zhao, Xiangyu Zhu et al.ICLR 2026 · 7 citations
