Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
Noy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope, Noam Slonim
Abstract
We introduce Debate Speech Evaluation as a novel and challenging benchmark for assessing LLM judges. Evaluating debate speeches requires a deep understanding of the speech at multiple levels, including argument strength and relevance, the coherence and organization of the speech, the appropriateness of its style and tone, and so on. This task involves a unique set of cognitive abilities that previously received limited attention in systematic LLM benchmarking. To explore such skills, we leverage a dataset of over 600 meticulously annotated debate speeches and present the first in-depth analysis of how state-of-the-art LLMs compare to human judges on this task. Our findings reveal a nuanced picture: while larger models can approximate individual human judgments in some respects, they differ substantially in their overall judgment behavior. We also investigate the ability of frontier LLMs to generate persuasive, opinionated speeches, showing that models may perform at a human level on this task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 846a2ac3-20e5-479e-bb93-12598f7683f8Cited by top-tier papers1
Ask how each one uses itBuilds on11
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
- A Large-Scale Dataset for Argument Quality Ranking: Construction and AnalysisShai Gretz, Roni Friedman, Edo Cohen-Karlik, Assaf Toledo et al.AAAI 2020 · 148 citations
- Corpus Wide Argument Mining - A Working SolutionLiat Ein-Dor, Eyal Shnarch, Lena Dankin, Alon Halfon et al.AAAI 2020 · 70 citations
- Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language ModelsSeungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin et al.EMNLP 2024 · 38 citations
Related papers
- ArgGenBench: Benchmarking the Complex Controlled Argument Generation Capability of Large Language ModelsBojun Jin, Jianzhu Bao, Yang Sun, Yice Zhang et al.ACL 2026
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality EvaluationHui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu et al.ACL 2026 · 21 citations
- Exploring the Potential of Large Language Models in Computational ArgumentationGuizhen Chen, Liying Cheng, Anh Tuan Luu, Lidong BingACL 2024 · 8 citations
- Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical FallaciesXudong Shen, Li Yuan, Ye Chen, Xin Wu et al.ACL 2026
- Debating with More Persuasive LLMs Leads to More Truthful AnswersAkbir Khan, John Hughes, Dan Valentine, Laura Ruis et al.ICML 2024 · 244 citations
