A Diagnostic Study of Multi-Agent LLMs for Real-World Debates
Priya Pitre, Gaurav Srivastava, Lu Zhang, Le Wang, Naren Ramakrishnan, Xuan Wang
Abstract
Multi-agent LLM debates are increasingly used in domains such as policy, politics, and city planning, where ground truth is often unavailable. Yet existing evaluations rely heavily on outcomebased proxies such as consensus, majority vote, or LLM-as-judge scores, which can miss failures like sycophancy, domination, and premature convergence. We introduce a diagnostic framework that evaluates both debate outcomes and the deliberative process using interpretable metrics for engagement, responsiveness, influence asymmetry, balance, stability, and agent utility. Across real-world debate settings and validation benchmarks, our process-level diagnostics align more closely with human judgments and reveal interaction failures that standard outcome-only measures overlook. These results show that reliable evaluation of multiagent debates requires measuring not only what answer agents reach, but how they reach it. Code is available at https://github.com/ priyapitre/DiagnosticStudyofLLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c42b794d-d5e4-47af-b97e-cca55f99f0f2Builds on15
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 651 citations
Related papers
- The Value of Variance: Mitigating Debate Collapse in Multi-Agent Systems via Uncertainty-Driven Policy OptimizationLuoxi Tang, Yuqiao Meng, Joseph Costa, Yingxue Zhang et al.ICML 2026 · 4 citations
- Multi-LLM Debate: Framework, Principals, and InterventionsAndrew Estornell, Yang LiuNeurIPS 2024 · 131 citations
- CIG: Measuring Conversational Information Gain in Deliberative Dialogues with Semantic Memory DynamicsMing-Bin Chen, Jey Han Lau, Lea FrermannACL 2026
- Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMsAndries P. Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D. Barrett et al.ICML 2024 · 82 citations
- Multi-Agent Debate for LLM Judges with Adaptive Stability DetectionTianyu Hu, Zhen Tan, Song Wang, Huaizhi Qu et al.NeurIPS 2025 · 25 citations
