Simple Agents, Biased Judges: Efficient Multi-Party Dialogue Generation & The Evaluation Gap
Kunal Samanta, Faisal Tareque Shohan, Amine Trabelsi, Richard Khoury
Abstract
Multi-party social dialogue remains underexplored in the literature, in part due to the difficulty and cost of evaluation. As a result, recent work on synthetic dialogue generation often relies on automated metrics and LLM-as-a-Judge frameworks, despite limited evidence that such judges reflect human preferences in social settings. In this work, we introduce a lightweight and controllable multi-party dialogue generation framework (MPOD) as an experimental instrument for studying generation and evaluation in social interaction. Using this framework, we conduct human evaluations of opendomain multi-party dialogue simulation and directly compare human judgments against stateof-the-art LLM judges. Across 319 pairwise comparisons, we observe near-random agreement between humans and automated judges (Cohen's κ ≈ 0.11), driven by systematic behaviors including extreme tie aversion and strong sensitivity to assistant-style verbosity. Crucially, human-human inter-annotator agreement (κ = 0.29) is substantially higher than human-LLM agreement. To isolate the mechanism underlying this misalignment, we introduce a controlled Transplant Ablation, showing that LLM judges consistently prefer conversations containing a single proprietary, assistantstyle agent. Additional stress tests show that judges prefer GPT-style conversations even when utterance order is randomly shuffled, indicating insensitivity to conversational structure and coherence. Our findings provide controlled evidence that current instruction-tuned LLM judges do not reliably reflect human preferences for naturalness, engagingness, and overall quality in multi-party social dialogue, calling into question their widespread use for validating synthetic conversational data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- Beyond the Surface: Measuring Self-Preference in LLM JudgmentsZhi-Yuan Chen, Hao Wang, Xinyu Zhang, Enrui Hu et al.EMNLP 2025
- Justice or Prejudice? Quantifying Biases in LLM-as-a-JudgeJiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen et al.ICLR 2025
Related papers
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang et al.ICML 2024 · 345 citations
- Don't Stop the Multi-Party! On Generating Synthetic Written Multi-Party Conversations with ConstraintsNicolò Penzo, Marco Guerini, Bruno Lepri, Goran Glavas et al.AAAI 2026 · 3 citations
- Bridging Human and LLM Judgments: Understanding and Narrowing the GapFelipe Maia Polo, Xinhe Wang, Mikhail Yurochkin, Gongjun Xu et al.NeurIPS 2025 · 7 citations
- Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationJiaju Chen, Yuxuan Lu, Xiaojie Wang, Huimin Zeng et al.ACL 2026 · 30 citations
- JuStRank: Benchmarking LLM Judges for System RankingAriel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim et al.ACL 2025
