ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
Siddhant Waghjale, Vishruth Veerendranath, Zhiruo Wang, Daniel Fried
摘要
Although large language models (LLMs) have been largely successful in generating functionally correct programs, conditioning models to produce efficient solutions while ensuring correctness remains a challenge. Further, unreliability in benchmarking code efficiency is a hurdle across varying hardware specifications for popular interpreted languages such as Python. In this paper, we present ECCO, a reproducible benchmark for evaluating program efficiency via two paradigms: natural language (NL) based code generation and history-based code editing. On ECCO, we adapt and thoroughly investigate the three most promising existing LLM-based approaches: in-context learning, iterative refinement with execution or NL feedback, and fine-tuning conditioned on execution and editing history. While most methods degrade functional correctness and moderately increase program efficiency, we find that adding execution information often helps maintain functional correctness, and NL feedback enhances more on efficiency. We release our benchmark to support future work on LLM-based generation of efficient code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Kevin: Multi-Turn RL for Generating CUDA KernelsCarlo Baronio, Pietro Marsella, Ben Pan, Simon Guo 等ICLR 2026 · 被引用 81 次
- Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency OptimizationMingzhe Du, Anh Tuan Luu, Yue Liu, Yuhao Qing 等NeurIPS 2025 · 被引用 18 次
- QuArch: A Benchmark for Evaluating LLM Reasoning in Computer ArchitectureShvetank Prakash, Andrew Cheng, Mark Mazumder, Arya Tschand 等ICML 2026 · 被引用 3 次
- FormulaCode: Evaluating Agentic Optimization on Large CodebasesAtharva Sehgal, James Hou, Akanksha Sarkar, Ishaan Mantripragada 等ICML 2026 · 被引用 3 次
- Speed Up Your Code: Progressive Code Acceleration Through Bidirectional Tree EditingLonghui Zhang, Jiahao Wang, Meishan Zhang, GaoXiong Cao 等ACL 2025 · 被引用 1 次
它引用的顶会 Paper9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- Generating Energy-Efficient Code via Large-Language Models - Where are we now?Radu Apsan, Vincenzo Stoico, Michel Albonico, Rudra Dhar 等ICSE 2026
- COFFE: A Code Efficiency Benchmark for Code GenerationYun Peng, Jun Wan, Yichen Li, Xiaoxue RenFSE 2025 · 被引用 8 次
- TRACE: Evaluating Execution Efficiency of LLM-Based Code TranslationZhihao Gong, Zeyu Sun, Dong Huang, Qingyuan Liang 等ACL 2026 · 被引用 5 次
- How efficient is LLM-generated code? A rigorous & high-standard benchmarkRuizhong Qiu, Weiliang Will Zeng, James Ezick, Christopher Lott 等ICLR 2025
- More Than Just Functional: LLM-as-a-Critique for Efficient Code GenerationDerui Zhu, Dingfan Chen, Jinfu Chen, Jens Grossklags 等NeurIPS 2025 · 被引用 2 次
