Lune

ICML2026Top-tier venue

A Diagnostic Study of Multi-Agent LLMs for Real-World Debates

Priya Pitre, Gaurav Srivastava, Lu Zhang, Le Wang, Naren Ramakrishnan, Xuan Wang

2026Year

Abstract

Multi-agent LLM debates are increasingly used in domains such as policy, politics, and city planning, where ground truth is often unavailable. Yet existing evaluations rely heavily on outcomebased proxies such as consensus, majority vote, or LLM-as-judge scores, which can miss failures like sycophancy, domination, and premature convergence. We introduce a diagnostic framework that evaluates both debate outcomes and the deliberative process using interpretable metrics for engagement, responsiveness, influence asymmetry, balance, stability, and agent utility. Across real-world debate settings and validation benchmarks, our process-level diagnostics align more closely with human judgments and reveal interaction failures that standard outcome-only measures overlook. These results show that reliable evaluation of multiagent debates requires measuring not only what answer agents reach, but how they reach it. Code is available at https://github.com/ priyapitre/DiagnosticStudyofLLMs.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c42b794d-d5e4-47af-b97e-cca55f99f0f2

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines