JudgeBench: A Benchmark for Evaluating LLM-Based Judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Yuan Tang, Alejandro Cuadron, Chenguang Wang, Raluca A. Popa, Ion Stoica
摘要
LLM-based judges have emerged as a scalable alternative to human evaluation and are increasingly used to assess, compare, and improve models. However, the reliability of LLM-based judges themselves is rarely scrutinized. As LLMs become more advanced, their responses grow more sophisticated, requiring stronger judges to evaluate them. Existing benchmarks primarily focus on a judge's alignment with human preferences, but often fail to account for more challenging tasks where crowdsourced human preference is a poor indicator of factual and logical correctness. To address this, we propose a novel evaluation framework to objectively evaluate LLM-based judges. Based on this framework, we propose JudgeBench, a benchmark for evaluating LLM-based judges on challenging response pairs spanning knowledge, reasoning, math, and coding. JudgeBench leverages a novel pipeline for converting existing difficult datasets into challenging response pairs with preference labels reflecting objective correctness. Our comprehensive evaluation on a collection of prompted judges, fine-tuned judges, multi-agent judges, and reward models shows that JudgeBench poses a significantly greater challenge than previous benchmarks, with many strong models (e.g., GPT-4o) performing just slightly better than random guessing. Overall, JudgeBench offers a reliable platform for assessing increasingly advanced LLM-based judges. Data and code are available at https://github.com/ ScalerLab/JudgeBench .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement LearningChenxi Whitehouse, Tianlu Wang, Ping Yu, Xian Li 等ICLR 2026 · 被引用 74 次
- Multi-Agent Debate for LLM Judges with Adaptive Stability DetectionTianyu Hu, Zhen Tan, Song Wang, Huaizhi Qu 等NeurIPS 2025 · 被引用 25 次
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement LearningRan Xu, Jingjing Chen, Jiayu Ye, Yu Wu 等ICLR 2026 · 被引用 17 次
- WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development QualityChunyang Li, Yilun Zheng, Xinting Huang, Tianqing Fang 等ICLR 2026 · 被引用 14 次
- EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?Zheng Lian, Licai Sun, Lan Chen, Haoyu Chen 等ICLR 2026 · 被引用 10 次
它引用的顶会 Paper15
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- Prometheus: Inducing Fine-Grained Evaluation Capability in Language ModelsSeungone Kim, Jamin Shin, Yejin Choi, Joel Jang 等ICLR 2024 · 被引用 468 次
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng 等ICLR 2024 · 被引用 368 次
相关 Paper
- ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for Mllm-Based Process JudgesJiaxin Ai, Pengfei Zhou, Zhaopan Xu, Ming Li 等ICCV 2025 · 被引用 9 次
- CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding TasksHongchao Jiang, Yiming Chen, Yushi Cao, Hung-Yi Lee 等ACL 2026 · 被引用 33 次
- Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual SettingsAustin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz 等ACL 2025 · 被引用 17 次
- Direct Judgement Preference OptimizationPeifeng Wang, Austin Xu, Yilun Zhou, Caiming Xiong 等EMNLP 2025 · 被引用 1 次
- WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation MetricsChenxu Liu, Yingjie Fu, Wei Yang, Ying Zhang 等ACL 2026 · 被引用 10 次
