COFFE: A Code Efficiency Benchmark for Code Generation
Yun Peng, Jun Wan, Yichen Li, Xiaoxue Ren
Abstract
Code generation has largely improved development efficiency in the era of large language models (LLMs). With the ability to follow instructions, current LLMs can be prompted to generate code solutions given detailed descriptions in natural language. Many research efforts are being devoted to improving the correctness of LLM-generated code, and many benchmarks are proposed to evaluate the correctness comprehensively. Despite the focus on correctness, the time efficiency of LLM-generated code solutions is under-explored. Current correctness benchmarks are not suitable for time efficiency evaluation since their test cases cannot well distinguish the time efficiency of different code solutions. Besides, the current execution time measurement is not stable and comprehensive, threatening the validity of the time efficiency evaluation.
To address the challenges in the time efficiency evaluation of code generation, we propose COFFE, a code generation benchmark for evaluating the time efficiency of LLM-generated code solutions. COFFE contains 398 and 358 problems for function-level and file-level code generation, respectively. To improve the distinguishability, we design a novel stressful test case generation approach with contracts and two new formats of test cases to improve the accuracy of generation. For the time evaluation metric, we propose efficienct@k based on CPU instruction count to ensure a stable and solid comparison between different solutions. We evaluate 14 popular LLMs on COFFE and identify four findings. Based on the findings, we draw some implications for LLM researchers and software practitioners to facilitate future research and usage of LLMs in code generation.
CCS Concepts: • Software and its engineering → Automatic programming.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- TRACE: Evaluating Execution Efficiency of LLM-Based Code TranslationZhihao Gong, Zeyu Sun, Dong Huang, Qingyuan Liang et al.ACL 2026 · 5 citations
- PEACE: Towards Efficient Project-Level Efficiency Optimization via Hybrid Code EditingXiaoxue Ren, Jun Wan, Yun Peng, Zhongxin Liu et al.ASE 2025 · 5 citations
- FormulaCode: Evaluating Agentic Optimization on Large CodebasesAtharva Sehgal, James Hou, Akanksha Sarkar, Ishaan Mantripragada et al.ICML 2026 · 3 citations
- More Than Just Functional: LLM-as-a-Critique for Efficient Code GenerationDerui Zhu, Dingfan Chen, Jinfu Chen, Jens Grossklags et al.NeurIPS 2025 · 2 citations
- Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and ReasoningXin Wang, Zhenhao Li, Zishuo DingASE 2025 · 1 citation
Builds on31
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
Related papers
- BEST: Benchmarking Efficiency in Space and Time for LLM-Generated CodeAocheng Shen, Boyu Zhang, Jiaze Li, Ruixuan Ma et al.ICML 2026
- How efficient is LLM-generated code? A rigorous & high-standard benchmarkRuizhong Qiu, Weiliang Will Zeng, James Ezick, Christopher Lott et al.ICLR 2025
- ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?Siddhant Waghjale, Vishruth Veerendranath, Zhiruo Wang, Daniel FriedEMNLP 2024 · 3 citations
- AutoCodeBench: Large Language Models are Automatic Code Benchmark GeneratorsChangzhi Zhou, Ao Liu, Yuchi Deng, Zhiying Zeng et al.ICLR 2026 · 27 citations
- Cracking Query Bottlenecks: Towards Efficiency-Oriented Text-to-SQL GenerationLi Lin, Yunfeng Shen, Lingfeng Bao, Rongxin Wu et al.ISSTA 2026
