How efficient is LLM-generated code? A rigorous & high-standard benchmark
Ruizhong Qiu, Weiliang Will Zeng, James Ezick, Christopher Lott, Hanghang Tong
Abstract
The emergence of large language models (LLMs) has significantly pushed the frontiers of program synthesis. Advancement of LLM-based program synthesis calls for a thorough evaluation of LLM-generated code. Most evaluation frameworks focus on the (functional) correctness of generated code; efficiency, as an important measure of code quality, has been overlooked in existing evaluations. In this work, we develop ENAMEL (EfficeNcy AutoMatic EvaLuator), a rigorous and high-standard benchmark for evaluating the capability of LLMs in generating efficient code. Firstly, we propose a new efficiency metric called eff@k, which generalizes the pass@k metric from correctness to efficiency and appropriately handles right-censored execution time. Furthermore, we derive an unbiased and variance-reduced estimator of eff@k via Rao-Blackwellization; we also provide a numerically stable implementation for the new estimator. Secondly, to set a high standard for efficiency evaluation, we employ a human expert to design best algorithms and implementations as our reference solutions of efficiency, many of which are much more efficient than existing canonical solutions in HumanEval and HumanEval+. Moreover, to ensure a rigorous evaluation, we employ a human expert to curate strong test case generators to filter out wrong code and differentiate suboptimal algorithms. An extensive study across 30 popular LLMs using our benchmark ENAMEL shows that LLMs still fall short of generating expert-level efficient code. Using two subsets of our problem set, we demonstrate that such deficiency is because current LLMs struggle in designing advanced algorithms and are barely aware of implementation optimization. Our benchmark is publicly available at https://github.com/q-rz/enamel.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5248ab2b-f5f0-4aa2-917a-329fa46ec6a2Cited by top-tier papers17
- EffiLearner: Enhancing Efficiency of Generated Code via Self-OptimizationDong Huang, Jianbo Dai, Han Weng, Puzhen Wu et al.NeurIPS 2024 · 54 citations
- ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMsJiaru Zou, Ling Yang, Jingwen Gu, Jiahao Qiu et al.NeurIPS 2025 · 51 citations
- Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency OptimizationMingzhe Du, Anh Tuan Luu, Yue Liu, Yuhao Qing et al.NeurIPS 2025 · 18 citations
- Gradient Compressed Sensing: A Query-Efficient Gradient Estimator for High-Dimensional Zeroth-Order OptimizationRuizhong Qiu, Hanghang TongICML 2024 · 12 citations
- Continual Low-Rank Adapters for LLM-based Generative Recommender SystemsHyunsik Yoo, Ting-Wei Li, SeongKu Kang, Zhining Liu et al.ICLR 2026 · 9 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
Related papers
- BEST: Benchmarking Efficiency in Space and Time for LLM-Generated CodeAocheng Shen, Boyu Zhang, Jiaze Li, Ruixuan Ma et al.ICML 2026
- TRACE: Evaluating Execution Efficiency of LLM-Based Code TranslationZhihao Gong, Zeyu Sun, Dong Huang, Qingyuan Liang et al.ACL 2026 · 5 citations
- COFFE: A Code Efficiency Benchmark for Code GenerationYun Peng, Jun Wan, Yichen Li, Xiaoxue RenFSE 2025 · 8 citations
- ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?Siddhant Waghjale, Vishruth Veerendranath, Zhiruo Wang, Daniel FriedEMNLP 2024 · 3 citations
- Self-Edit: Fault-Aware Code Editor for Code GenerationKechi Zhang, Zhuo Li, Jia Li, Ge Li et al.ACL 2023 · 42 citations
