CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Hongchao Jiang, Yiming Chen, Yushi Cao, Hung-Yi Lee, Robby T. Tan
摘要
Large Language Models (LLMs) have significantly advanced the state-of-the-art in various coding tasks. Beyond directly answering user queries, LLMs can also serve as judges, assessing and comparing the quality of responses generated by other models. Such an evaluation capability is crucial both for benchmarking different LLMs and for improving response quality through response ranking. However, despite the growing adoption of the LLM-as-a-Judge paradigm, its effectiveness in coding scenarios remains underexplored due to the absence of dedicated benchmarks. To address this gap, we introduce CodeJudgeBench, a benchmark explicitly designed to evaluate the performance of LLM-as-a-Judge models across three critical coding tasks: code generation, code repair, and unit test generation. Through comprehensive benchmarking of 26 LLM-as-a-Judge models, we find that recent thinking models significantly outperform non-thinking models on our carefully designed code judging tasks. Notably, even relatively small thinking models, such as Qwen3-8B, can outperform specially trained LLM-as-a-Judge models up to 70B in size. Nevertheless, all models still exhibit significant randomness in their judgment of coding tasks. For pairwise judging tasks, simply changing the order in which responses are presented can substantially impact accuracy. In addition, when judging code and unit tests written by different LLMs, LLM-as-a-Judge models also show variance in performance. This sensitivity raises concerns about the reliability and consistency of LLM-as-a-Judge in coding scenarios. Lastly, we study optimal prompting strategies for LLM-as-a-Judge. We find that using pair-wise comparison outperforms scalar point-wise judging. Furthermore, retaining comments and reasoning in the full, unprocessed LLM response leads to improved judge performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement LearningRan Xu, Jingjing Chen, Jiayu Ye, Yu Wu 等ICLR 2026 · 被引用 17 次
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric DomainsAustin Xu, Xuan-Phi Nguyen, Yilun Zhou, Chien-Sheng Wu 等ICLR 2026 · 被引用 8 次
- Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang 等EMNLP 2025 · 被引用 2 次
- Powering Verifiable Learning via Automated Evolutionary Data SynthesisHe Du, Bowen Li, Aijun Yang, Siyang He 等ACL 2026
它引用的顶会 Paper18
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng 等ICLR 2024 · 被引用 368 次
- Generative Judge for Evaluating AlignmentJunlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan 等ICLR 2024 · 被引用 173 次
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 被引用 172 次
- RM-R1: Reward Modeling as ReasoningXiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin 等ICLR 2026 · 被引用 147 次
相关 Paper
- CodeJudge: Evaluating Code Generation with Large Language ModelsWeixi Tong, Tianyi ZhangEMNLP 2024 · 被引用 25 次
- Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling EvaluatorsYilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong 等ICML 2025
- Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive ProgrammingTingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu 等ICML 2026 · 被引用 1 次
- QuanBench: Benchmarking Quantum Code Generation with Large Language ModelsXiaoyu Guo, Minggu Wang, Jianjun ZhaoASE 2025 · 被引用 5 次
- DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code GenerationQiming Zhu, Jialun Cao, Yaojie Lu, Hongyu Lin 等AAAI 2025 · 被引用 25 次
