E2EDev: Benchmarking Large Language Models in End-to-End Software Development Task
Jingyao Liu, Chen Huang, Zhizhao Guan, Wenqiang Lei, Yang Deng
Abstract
The rapid advancement in large language models (LLMs) has demonstrated significant potential in End-to-End Software Development (E2ESD). However, existing E2ESD benchmarks are limited by coarse-grained requirement specifications and unreliable evaluation protocols, hindering a true understanding of current framework capabilities. To address these limitations, we present E2EDev, a novel benchmark grounded in the principles of Behavior-Driven Development (BDD), which evaluates the capabilities of E2ESD frameworks by assessing whether the generated software meets user needs through mimicking real user interactions (Figure 1). E2EDev comprises (i) a fine-grained set of user requirements, (ii) multiple BDD test scenarios with corresponding Python step implementations for each requirement, and (iii) a fully automated testing pipeline built on the Behave framework. To ensure its quality while reducing the annotation effort, E2EDev leverages our proposed Human-in-the-Loop Multi-Agent Annotation Framework (HITL-MAA). By evaluating various E2ESD frameworks and LLM backbones with E2EDev, our analysis reveals a persistent struggle to effectively solve these tasks, underscoring the critical need for more effective and cost-efficient E2ESD solutions. Our codebase and benchmark are publicly available at https://github.com/SCUNLP/E2EDev.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97c62500-8d1d-4fdb-9e9c-7022207d535fBuilds on5
- Large Language Models as Analogical ReasonersMichihiro Yasunaga, Xinyun Chen, Yujia Li, Panupong Pasupat et al.ICLR 2024 · 155 citations
- LLMCarbon: Modeling the End-to-End Carbon Footprint of Large Language ModelsAhmad Faiz, Sotaro Kaneda, Ruhan Wang, Rita Chukwunyere Osi et al.ICLR 2024 · 129 citations
- CoderEval: A Benchmark of Pragmatic Code Generation with Generative Pre-trained ModelsHao Yu, Bo Shen, Dezhi Ran, Jiaxin Zhang et al.ICSE 2024 · 107 citations
- ELABORATION: A Comprehensive Benchmark on Human-LLM Competitive ProgrammingXinwei Yang, Zhaofeng Liu, Chen Huang, Jiashuai Zhang et al.ACL 2025 · 15 citations
- ChatDev: Communicative Agents for Software DevelopmentChen Qian, Wei Liu, Hongzhang Liu, Nuo Chen et al.ACL 2024
Related papers
- Towards Iterative End-to-End Software Development: A Feature-Driven Multi-agent FrameworkJunwei Liu, Chen Xu, Chong Wang, Tong Bai et al.ISSTA 2026 · 1 citation
- Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering TasksDimitrios Rontogiannis, Maxime Peyrard, Nicolas Mario Baldwin, Martin Josifoski et al.AAAI 2026 · 1 citation
- Feature-Driven End-to-End Test GenerationParsa Alian, Noor Nashid, Mobina Shahbandeh, Taha Shabani et al.ICSE 2025 · 2 citations
- TraceDev: A Traceability-Driven Multi-agent Framework for Requirement-to-Code DevelopmentMingyu Chen, Yakun Zhang, Zihao Xie, Yixing Luo et al.ISSTA 2026
- WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation MetricsChenxu Liu, Yingjie Fu, Wei Yang, Ying Zhang et al.ACL 2026 · 10 citations
