Simple Agents, Biased Judges: Efficient Multi-Party Dialogue Generation & The Evaluation Gap
Kunal Samanta, Faisal Tareque Shohan, Amine Trabelsi, Richard Khoury
摘要
Multi-party social dialogue remains underexplored in the literature, in part due to the difficulty and cost of evaluation. As a result, recent work on synthetic dialogue generation often relies on automated metrics and LLM-as-a-Judge frameworks, despite limited evidence that such judges reflect human preferences in social settings. In this work, we introduce a lightweight and controllable multi-party dialogue generation framework (MPOD) as an experimental instrument for studying generation and evaluation in social interaction. Using this framework, we conduct human evaluations of opendomain multi-party dialogue simulation and directly compare human judgments against stateof-the-art LLM judges. Across 319 pairwise comparisons, we observe near-random agreement between humans and automated judges (Cohen's κ ≈ 0.11), driven by systematic behaviors including extreme tie aversion and strong sensitivity to assistant-style verbosity. Crucially, human-human inter-annotator agreement (κ = 0.29) is substantially higher than human-LLM agreement. To isolate the mechanism underlying this misalignment, we introduce a controlled Transplant Ablation, showing that LLM judges consistently prefer conversations containing a single proprietary, assistantstyle agent. Additional stress tests show that judges prefer GPT-style conversations even when utterance order is randomly shuffled, indicating insensitivity to conversational structure and coherence. Our findings provide controlled evidence that current instruction-tuned LLM judges do not reliably reflect human preferences for naturalness, engagingness, and overall quality in multi-party social dialogue, calling into question their widespread use for validating synthetic conversational data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- Beyond the Surface: Measuring Self-Preference in LLM JudgmentsZhi-Yuan Chen, Hao Wang, Xinyu Zhang, Enrui Hu 等EMNLP 2025
- Justice or Prejudice? Quantifying Biases in LLM-as-a-JudgeJiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen 等ICLR 2025
相关 Paper
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang 等ICML 2024 · 被引用 345 次
- Don't Stop the Multi-Party! On Generating Synthetic Written Multi-Party Conversations with ConstraintsNicolò Penzo, Marco Guerini, Bruno Lepri, Goran Glavas 等AAAI 2026 · 被引用 3 次
- Bridging Human and LLM Judgments: Understanding and Narrowing the GapFelipe Maia Polo, Xinhe Wang, Mikhail Yurochkin, Gongjun Xu 等NeurIPS 2025 · 被引用 7 次
- Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationJiaju Chen, Yuxuan Lu, Xiaojie Wang, Huimin Zeng 等ACL 2026 · 被引用 30 次
- JuStRank: Benchmarking LLM Judges for System RankingAriel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim 等ACL 2025
