AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Changzhi Zhou, Ao Liu, Yuchi Deng, Zhiying Zeng, Tao Zhang, Haotian Zhu, Jianwei Cai, Yue Mao, Chenchen Zhang, Lingyun Tan, Ziyan Xu, Bohui Zhai
Abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains, with code generation emerging as a key area of focus. While numerous benchmarks have been proposed to evaluate their code generation abilities, these benchmarks face several critical limitations. First, they often rely on manual annotations, which are time-consuming and difficult to scale across different programming languages and problem complexities. Second, most existing benchmarks focus primarily on Python, while the few multilingual benchmarks suffer from limited difficulty and uneven language distribution. To address these challenges, we propose AutoCodeGen, an automated method for generating high-difficulty multilingual code generation datasets without manual annotations. AutoCodeGen ensures the correctness and completeness of test cases by generating test inputs with LLMs and obtaining test outputs through a multilingual sandbox, while achieving high data quality through reverse-order problem generation and multiple filtering steps. Using this novel method, we introduce AutoCodeBench, a large-scale code generation benchmark comprising 3,920 problems evenly distributed across 20 programming languages. It is specifically designed to evaluate LLMs on challenging, diverse, and practical multilingual tasks. We evaluate over 30 leading open-source and proprietary LLMs on AutoCodeBench and its simplified version AutoCodeBench-Lite. The results show that even the most advanced LLMs struggle with the complexity, diversity, and multilingual nature of these tasks. Besides, we introduce AutoCodeBench-Complete, specifically designed for base models to assess their few-shot code generation capabilities. We hope the AutoCodeBench series will serve as a valuable resource and inspire the community to focus on more challenging and practical multilingual code generation scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c77918f-e696-48b1-8110-fed5d7466375Cited by top-tier papers7
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesMike A. Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li et al.ICLR 2026 · 520 citations
- Aegis: Automated Error Generation and Attribution for Multi-Agent SystemsFanqi Kong, Ruijie Zhang, Huaxiao Yin, Guibin Zhang et al.ICLR 2026 · 16 citations
- MLE-Smith: Scaling MLE Tasks with Automated Multi-agent PipelineRushi Qiang, Yuchen Zhuang, Anikait Singh, Percy Liang et al.ICLR 2026 · 7 citations
- Code2Bench: Scaling Source and Rigor for Dynamic Benchmark ConstructionZhe Zhang, Runlin Liu, Aishan Liu, Xingyu Liu et al.ICLR 2026 · 4 citations
- Beyond Text-to-SQL: Can LLMs Really Debug Enterprise ETL SQL?Jing Ye, Yiwen Duan, Yonghong Yu, Victor Ma et al.ICML 2026 · 1 citation
Builds on11
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- Magicoder: Empowering Code Generation with OSS-InstructYuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding et al.ICML 2024 · 246 citations
Related papers
- McEval: Massively Multilingual Code EvaluationLinzheng Chai, Shukai Liu, Jian Yang, Yuwei Yin et al.ICLR 2025 · 1 citation
- Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation BenchmarkDewu Zheng, Yanlin Wang, Ensheng Shi, Xilin Liu et al.ICSE 2026
- DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code GenerationQiming Zhu, Jialun Cao, Yaojie Lu, Hongyu Lin et al.AAAI 2025 · 25 citations
- Multi-LCB: Extending LiveCodeBench to Multiple Programming LanguagesMaria Ivanova, Pavel Zadorozhny, Rodion Levichev, Ivan Petrov et al.ICLR 2026 · 3 citations
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeNaman Jain, King Han, Alex Gu, Wen-Ding Li et al.ICLR 2025
