JuStRank: Benchmarking LLM Judges for System Ranking
Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim, Lilach Eden, Asaf Yehudai
摘要
Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and versatility of such evaluations make the use of LLMbased judges a compelling solution for this challenge. Crucially, this approach requires first to validate the quality of the LLM judge itself. Previous work has focused on instance-based assessment of LLM judges, where a judge is evaluated over a set of responses, or response pairs, while being agnostic to their source systems. We argue that this setting overlooks critical factors affecting system-level ranking, such as a judge's positive or negative bias towards certain systems. To address this gap, we conduct the first large-scale study of LLM judges as system rankers. System scores are generated by aggregating judgment scores over multiple system outputs, and the judge's quality is assessed by comparing the resulting system ranking to a human-based ranking. Beyond overall judge assessment, our analysis provides a fine-grained characterization of judge behavior, including their decisiveness and bias.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi 等EMNLP 2025 · 被引用 37 次
- On Evaluating LLM Alignment by Evaluating LLMs as JudgesYixin Liu, Pengfei Liu, Arman CohanNeurIPS 2025 · 被引用 7 次
- AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language ReasoningXiping Li, Jianghong MaACL 2026 · 被引用 3 次
- Debatable Intelligence: Benchmarking LLM Judges via Debate Speech EvaluationNoy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope 等EMNLP 2025 · 被引用 1 次
- Mediocrity is the key for LLM as a Judge Anchor SelectionShachar Don-Yehiya, Asaf Yehudai, Leshem Choshen, Omri AbendACL 2026 · 被引用 1 次
它引用的顶会 Paper12
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsRui Yang, Ruomeng Ding, Yong Lin, Huan Zhang 等NeurIPS 2024 · 被引用 157 次
- ODIN: Disentangled Reward Mitigates Hacking in RLHFLichang Chen, Chen Zhu, Jiuhai Chen, Davit Soselia 等ICML 2024 · 被引用 119 次
- Achieving Human Parity in Content-Grounded Datasets GenerationAsaf Yehudai, Boaz Carmeli, Yosi Mass, Ofir Arviv 等ICLR 2024 · 被引用 9 次
- Label-Efficient Model Selection for Text GenerationShir Ashury-Tahan, Ariel Gera, Benjamin Sznajder, Leshem Choshen 等ACL 2024 · 被引用 1 次
相关 Paper
- Validating LLM-as-a-Judge Systems under Rating IndeterminacyLuke Guerdan, Solon Barocas, Kenneth Holstein, Hanna M. Wallach 等NeurIPS 2025 · 被引用 31 次
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang 等ICML 2024 · 被引用 345 次
- Beyond the Surface: Measuring Self-Preference in LLM JudgmentsZhi-Yuan Chen, Hao Wang, Xinyu Zhang, Enrui Hu 等EMNLP 2025
- Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference EvaluationsDani Roytburg, Matthew Bozoukov, Matthew Nguyen, Jou Barzdukas 等ICML 2026
- Bridging Human and LLM Judgments: Understanding and Narrowing the GapFelipe Maia Polo, Xinhe Wang, Mikhail Yurochkin, Gongjun Xu 等NeurIPS 2025 · 被引用 7 次
