Lune

ICML2026顶会

BEST: Benchmarking Efficiency in Space and Time for LLM-Generated Code

Aocheng Shen, Boyu Zhang, Jiaze Li, Ruixuan Ma, Qiankun Zhang, Wang, Bin Yuan, Shenghao Liu, Xianjun Deng

出版方
2026年份

摘要

Large language models (LLMs) have revolutionized research in software engineering, and among various tasks, LLM-based code synthesis is promising. A recent line of benchmarks aims to evaluate LLM-generated codes in time efficiency, beyond their correctness. However, space, another vital aspect of code efficiency, is rarely evaluated in prior benchmarks. To fill in the gap, this paper introduces BEST, the first benchmark for evaluating the efficiency of LLM-generated codes in both time and space. It comprises 440440 coding tasks that are rigorously constructed by experts. In addition, we propose a fine-grained subtask-based evaluation scheme by dividing each task into multiple subtasks, with different input scales and difficulties. Each subtask is then accompanied by an expert-crafted standard implementation as the efficiency baseline, which achieves the Pareto optimum. Building on BEST, we introduce a unified and novel dual-indicator (time and space) metric, named dual@k{k}, generalizing the notion of the standard pass@k{k} metric and building on a careful and novel construction of a weight matrix of subtasks. Through extensive experiments with dual@k{k} across 5050 LLMs on BEST, our evaluation demonstrates that while LLMs exhibit weak capabilities in generating time-efficient code, their capabilities in space-efficient code generation are even worse. The benchmark is provided at https://github.com/kmsgk0/BEST.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 30a087e0-9bc2-479b-9072-8e068a633f97

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖