ACL2026

Simple Agents, Biased Judges: Efficient Multi-Party Dialogue Generation & The Evaluation Gap

Kunal Samanta, Faisal Tareque Shohan, Amine Trabelsi, Richard Khoury

摘要

Multi-party social dialogue remains underexplored in the literature, in part due to the difficulty and cost of evaluation. As a result, recent work on synthetic dialogue generation often relies on automated metrics and LLM-as-a-Judge frameworks, despite limited evidence that such judges reflect human preferences in social settings. In this work, we introduce a lightweight and controllable multi-party dialogue generation framework (MPOD) as an experimental instrument for studying generation and evaluation in social interaction. Using this framework, we conduct human evaluations of opendomain multi-party dialogue simulation and directly compare human judgments against stateof-the-art LLM judges. Across 319 pairwise comparisons, we observe near-random agreement between humans and automated judges (Cohen's κ ≈ 0.11), driven by systematic behaviors including extreme tie aversion and strong sensitivity to assistant-style verbosity. Crucially, human-human inter-annotator agreement (κ = 0.29) is substantially higher than human-LLM agreement. To isolate the mechanism underlying this misalignment, we introduce a controlled Transplant Ablation, showing that LLM judges consistently prefer conversations containing a single proprietary, assistantstyle agent. Additional stress tests show that judges prefer GPT-style conversations even when utterance order is randomly shuffled, indicating insensitivity to conversational structure and coherence. Our findings provide controlled evidence that current instruction-tuned LLM judges do not reliably reflect human preferences for naturalness, engagingness, and overall quality in multi-party social dialogue, calling into question their widespread use for validating synthetic conversational data.