Investigating Non-Transitivity in LLM-as-a-Judge
Yi Xu, Laura Ruis, Tim Rocktäschel, Robert Kirk
摘要
Automatic evaluation methods based on large language models (LLMs) are emerging as the standard tool for assessing the instruction-following abilities of LLM-based agents. The most common method in this paradigm, pairwise comparisons with a baseline model, critically depends on the assumption of transitive preferences. However, the validity of this assumption remains largely unexplored. In this study, we investigate the presence of non-transitivity within the AlpacaEval framework and analyze its effects on model rankings. We find that LLM judges exhibit non-transitive preferences, leading to rankings that are sensitive to the choice of the baseline model. To mitigate this issue, we show that round-robin tournaments combined with Bradley-Terry models of preference can produce more reliable rankings. Notably, our method increases both the Spearman correlation and the Kendall correlation with Chatbot Arena (95.0% → 96.4% and 82.1% → 86.3% respectively). To address the computational cost of round-robin tournaments, we propose Swiss-Wise Iterative Matchmaking (SWIM) tournaments, using a dynamic matching strategy to capture the benefits of round-robin tournaments while maintaining computational efficiency. Investigating Non-Transitivity in LLM-as-a-Judge Judge Evaluation A B C Construct preference matrices for each instruction. Conduct pairwise comparisons through a round robin tournament. Compute Elo scores with the Bradley-Terry model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate ThemYidong Wang, Yunze Song, Tingyuan Zhu, Xuanwang Zhang 等ICLR 2026 · 被引用 29 次
- Generative Value Conflicts Reveal LLM PrioritiesAndy Liu, Kshitish Ghate, Mona T. Diab, Daniel Fried 等ICLR 2026 · 被引用 17 次
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential GameBarna Pásztor, Thomas Kleine Buening, Andreas KrauseICLR 2026 · 被引用 9 次
- Gained in Translation: Privileged Pairwise Judges Enhance Multilingual ReasoningLintang Sutawika, Gokul Swamy, Steven Wu, Graham NeubigACL 2026 · 被引用 6 次
- Nonparametric LLM Evaluation from Preference DataDennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan FeuerriegelICML 2026 · 被引用 5 次
它引用的顶会 Paper13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Debating with More Persuasive LLMs Leads to More Truthful AnswersAkbir Khan, John Hughes, Dan Valentine, Laura Ruis 等ICML 2024 · 被引用 244 次
- Rainbow Teaming: Open-Ended Generation of Diverse Adversarial PromptsMikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro 等NeurIPS 2024 · 被引用 231 次
- Real World Games Look Like Spinning TopsWojciech M. Czarnecki, Gauthier Gidel, Brendan D. Tracey, Karl Tuyls 等NeurIPS 2020 · 被引用 123 次
相关 Paper
- ELSPR: Evaluator LLM Training Data Self-Purification on Non-Transitive Preferences via Tournament Graph ReconstructionYan Yu, Yilun Liu, Minggui He, Shimin Tao 等AAAI 2026 · 被引用 2 次
- Dropping Just a Handful of Preferences Can Change Top Large Language Model RankingsJenny Y. Huang, Yunyi Shen, Dennis Wei, Tamara BroderickICLR 2026 · 被引用 8 次
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker 等NeurIPS 2024 · 被引用 94 次
- Prediction-Powered Ranking of Large Language ModelsIvi Chatzi, Eleni Straitouri, Suhas Thejaswi, Manuel Gomez RodriguezNeurIPS 2024 · 被引用 34 次
- Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee DiscussionsRuochen Zhao, Wenxuan Zhang, Yew Ken Chia, Weiwen Xu 等ACL 2025 · 被引用 34 次
