HumanEvo: An Evolution-Aware Benchmark for More Realistic Evaluation of Repository-Level Code Generation
Dewu Zheng, Yanlin Wang, Ensheng Shi, Ruikai Zhang, Yuchi Ma, Hongyu Zhang, Zibin Zheng
摘要
To evaluate the repository-level code generation capabilities of Large Language Models (LLMs) in complex real-world software development scenarios, many evaluation methods have been developed. These methods typically leverage contextual code from the latest version of a project to assist LLMs in accurately generating the desired function. However, such evaluation methods fail to consider the dynamic evolution of software projects over time, which we refer to as evolution-ignored settings. This in turn results in inaccurate evaluation of LLMs' performance. In this paper, we conduct an empirical study to deeply understand LLMs' code generation performance within settings that reflect the evolution nature of software development. To achieve this, we first construct an evolution-aware repository-level code generation dataset, namely HumanEvo, equipped with an automated execution-based evaluation tool. Second, we manually categorize HumanEvo according to dependency levels to more comprehensively analyze the model's performance in generating functions with different dependency levels. Third, we conduct extensive experiments on HumanEvo with seven representative and diverse LLMs to verify the effectiveness of the proposed benchmark. We obtain several important findings through our experimental study. For example, we find that previous evolution-ignored evaluation methods result in inflated performance of LLMs, with performance overestimations ranging from 10.0% to 61.1% under different context acquisition methods, compared to the evolution-aware evaluation approach. Based on the findings, we give actionable suggestions for more realistic evaluation of LLMs on code generation. We also build a shared evolution-aware code generation toolbox to facilitate future research. The replication package including source code and datasets is anonymously available at https://github.Com/DeepSoftwareAnalytics/HumanEvo.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- DrainCode: Stealthy Energy Consumption Attacks on Retrieval-Augmented Code Generation via Context PoisoningYanli Wang, Jiadong Wu, Tianyue Jiang, Mingwei Liu 等ASE 2025 · 被引用 3 次
- AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code CompletionTianyue Jiang, Yanlin Wang, Yanli Wang, Daya Guo 等ASE 2025 · 被引用 2 次
- Code-DiTing: Automatic Evaluation of Code Generation without References or Test CasesGuang Yang, Yu Zhou, Xiang Chen, Wei Zheng 等ASE 2025 · 被引用 1 次
- From Static to Dynamic: Benchmarking Real-World Code Review with MCR-BenchDewu Zheng, Yanlin Wang, Xiwen Wang, Kefeng Duan 等ISSTA 2026
- Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation BenchmarkDewu Zheng, Yanlin Wang, Ensheng Shi, Xilin Liu 等ICSE 2026
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang 等ICML 2023 · 被引用 504 次
相关 Paper
- RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development PracticesJia Li, Hongyi Deng, Yiran Zhang, Kechi Zhang 等FSE 2026
- Evaluating Large Language Models in Class-Level Code GenerationXueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang 等ICSE 2024 · 被引用 118 次
- Towards Understanding the Characteristics of Code Generation Errors Made by Large Language ModelsZhijie Wang, Zijie Zhou, Da Song, Yuheng Huang 等ICSE 2025 · 被引用 12 次
- DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code GenerationQiming Zhu, Jialun Cao, Yaojie Lu, Hongyu Lin 等AAAI 2025 · 被引用 25 次
- Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'Shanchao Liang, Nan Jiang, Yiran Hu, Lin TanACL 2025 · 被引用 9 次
