CodeJudge: Evaluating Code Generation with Large Language Models
Weixi Tong, Tianyi Zhang
Abstract
Large Language Models (LLMs) have shown promising performance in code generation. However, how to reliably evaluate code generated by LLMs remains an unresolved problem. This paper presents CODEJUDGE , a code evaluation framework that leverages LLMs to evaluate the semantic correctness of generated code without the need for test cases. We investigate different ways to guide the LLM in performing "slow thinking" to arrive at an in-depth and reliable evaluation. We experimented with four LLMs as evaluators on four code generation datasets and five programming languages. The results show that CODEJUDGE significantly outperformed existing methods in most settings. Furthermore, compared with a SOTA GPT-3.5-based code evaluation method, CODE-JUDGE achieved better results even when using a much smaller model, Llama-3-8B-Instruct. Our code and datasets are available on GitHub https://github.com/VichyTong/CodeJudge .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be223a72-0962-478f-a17e-d3cd5f25c9f7Cited by top-tier papers19
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- KRAMABENCH: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data LakesEugenie Lai, Gerardo Vitagliano, Ziyu Zhang, Om Chabra et al.ICLR 2026 · 37 citations
- Agents Under Siege: Breaking Pragmatic Multi-Agent LLM Systems with Optimized Prompt AttacksRana Muhammad Shahroz, Zhen Tan, Sukwon Yun, Charles Fleming et al.ACL 2025 · 18 citations
- Towards Understanding the Characteristics of Code Generation Errors Made by Large Language ModelsZhijie Wang, Zijie Zhou, Da Song, Yuheng Huang et al.ICSE 2025 · 12 citations
- Evaluating Text Creativity across Diverse Domains: a Dataset and Large Language Model EvaluatorQian Cao, Xiting Wang, Yuzhuo Yuan, Yahui Liu et al.ICLR 2026 · 11 citations
Builds on11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu et al.ICLR 2024 · 871 citations
Related papers
- CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding TasksHongchao Jiang, Yiming Chen, Yushi Cao, Hung-Yi Lee et al.ACL 2026 · 33 citations
- Can You Really Trust Code Copilot? Evaluating Large Language Models from a Code Security PerspectiveYutao Mou, Xiao Deng, Yuxiao Luo, Shikun Zhang et al.ACL 2025 · 4 citations
- Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChainMarcus J. Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar et al.ICLR 2024 · 39 citations
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationYuchen Tian, Weixiang Yan, Qian Yang, Xuandong Zhao et al.AAAI 2025 · 41 citations
- Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling EvaluatorsYilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong et al.ICML 2025
