FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
Wei Li, Xin Zhang, Zhongxin Guo, Shaoguang Mao, Wen Luo, Guangyue Peng, Yangyu Huang, Houfeng Wang, Scarlett Li
摘要
Implementing new features in repository-level codebases is a crucial application of code generation models. However, current benchmarks lack a dedicated evaluation framework for this capability. To fill this gap, we introduce FEA-Bench, a benchmark designed to assess the ability of large language models (LLMs) to perform incremental development within code repositories. We collect pull requests from 83 GitHub repositories and use rule-based and intent-based filtering to construct task instances focused on new feature development. Each task instance containing code changes is paired with relevant unit test files to ensure that the solution can be verified. The feature implementation requires LLMs to simultaneously possess code completion capabilities for new components and code editing abilities for other relevant parts in the code repository, providing a more comprehensive evaluation method of LLMs' automated software engineering capabilities. Experimental results show that LLMs perform significantly worse in the FEA-Bench, highlighting considerable challenges in such repository-level incremental code development. Our code will soon be publicly available at https://github.com/microsoft/FEA-Bench .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- RPG: A Repository Planning Graph for Unified and Scalable Codebase GenerationJane Luo, Xin Zhang, Steven Liu, Jie Wu 等ICLR 2026 · 被引用 18 次
- Structurally Aligned Subtask-Level Memory for Software Engineering AgentsKangning Shen, Jingyuan Zhang, Chenxi Sun, Wencong Zeng 等ICML 2026 · 被引用 9 次
- Issue Localization via LLM-Driven Iterative Code Graph SearchingZhonghao Jiang, Xiaoxue Ren, Meng Yan, Wei Jiang 等ASE 2025 · 被引用 6 次
- SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing ScenariosJunkai Chen, Huihui Huang, Yunbo Lyu, Junwen An 等ACL 2026 · 被引用 5 次
- RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to RepositoryZhiyuan Peng, Xin Yin, Pu Zhao, Fangkai Yang 等ACL 2026 · 被引用 4 次
它引用的顶会 Paper13
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- Magicoder: Empowering Code Generation with OSS-InstructYuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding 等ICML 2024 · 被引用 246 次
- Repair Is Nearly Generation: Multilingual Program Repair with LLMsHarshit Joshi, José Pablo Cambronero Sánchez, Sumit Gulwani, Vu Le 等AAAI 2023 · 被引用 182 次
相关 Paper
- SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?Xinyi He, Qian Liu, Mingzhe Du, Lin Yan 等ICML 2026 · 被引用 31 次
- RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development PracticesJia Li, Hongyi Deng, Yiran Zhang, Kechi Zhang 等FSE 2026
- Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'Shanchao Liang, Nan Jiang, Yiran Hu, Lin TanACL 2025 · 被引用 9 次
- FeatureBench: Benchmarking Agentic Coding for Complex Feature DevelopmentQixing Zhou, Jiacheng Zhang, Haiyang Wang, Rui Hao 等ICLR 2026 · 被引用 30 次
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment GenerationZhengran Zeng, Ruikai Shi, Keke Han, Yixin Li 等FSE 2026
