Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Yilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong, Shafiq Joty
Abstract
Scaling test-time computation, or affording a generator large language model (LLM) extra compute during inference, typically employs the help of external non-generative evaluators (i.e., reward models). Concurrently, LLM-judges, models trained to generate evaluations and critiques (explanations) in natural language, are becoming increasingly popular in automatic evaluation. Despite judge empirical successes, their effectiveness as evaluators in test-time scaling settings is largely unknown. In this paper, we introduce the Judge Evaluation for Test-Time Scaling (JETTS) benchmark, which evaluates judge performance in three domains (math reasoning, code generation, and instruction following) under three task settings: response reranking, step-level beam search, and critique-based response refinement. We evaluate 10 different judge models (7B-70B parameters) for 8 different base generator models (6.7B-72B parameters). Our benchmark shows that while judges are competitive with outcome reward models in reranking, they are consistently worse than process reward models in beam search procedures. Furthermore, though unique to LLM-judges, their natural language critiques are currently ineffective in guiding the generator towards better responses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e81b5d13-de6c-4f70-b320-fa7bf09db041Cited by top-tier papers11
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding TasksHongchao Jiang, Yiming Chen, Yushi Cao, Hung-Yi Lee et al.ACL 2026 · 33 citations
- Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect JudgesChen Feng, Minghe Shen, Ananth Balashankar, Carsten Gerner-Beuerle et al.ICLR 2026 · 24 citations
- Variation in Verification: Understanding Verification Dynamics in Large Language ModelsYefan Zhou, Austin Xu, Yilun Zhou, Janvijay Singh et al.ICLR 2026 · 17 citations
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement LearningRan Xu, Jingjing Chen, Jiayu Ye, Yu Wu et al.ICLR 2026 · 17 citations
Builds on8
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin et al.ACL 2025 · 209 citations
- Style Outweighs Substance: Failure Modes of LLM Judges in Alignment BenchmarkingBenjamin Feuer, Micah Goldblum, Teresa Datta, Sanjana Nambiar et al.ICLR 2025 · 1 citation
Related papers
- Think-J: Learning to Think for Generative LLM-as-a-JudgeHui Huang, Yancheng He, Hongli Zhou, Rui Zhang et al.AAAI 2026 · 12 citations
- Direct Judgement Preference OptimizationPeifeng Wang, Austin Xu, Yilun Zhou, Caiming Xiong et al.EMNLP 2025 · 1 citation
- CodeJudge: Evaluating Code Generation with Large Language ModelsWeixi Tong, Tianyi ZhangEMNLP 2024 · 25 citations
- VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video UnderstandingAbdul Waheed, Zhen Wu, Dareen Safar Alharthi, Seungone Kim et al.ICLR 2026 · 4 citations
- JudgeBench: A Benchmark for Evaluating LLM-Based JudgesSijun Tan, Siyuan Zhuang, Kyle Montgomery, William Yuan Tang et al.ICLR 2025
