TraceDev: A Traceability-Driven Multi-agent Framework for Requirement-to-Code Development
Mingyu Chen, Yakun Zhang, Zihao Xie, Yixing Luo, Jinrui Xu, Cuiyun Gao, Kaiqi Zhao, Yunming Ye
Abstract
In modern software development, the rapid advancement of Large Language Models (LLMs) has made the end-to-end transformation of Natural Language Requirements (NLRs) into executable repository-level code increasingly feasible. However, existing approaches typically rely on simplified instructions (e.g., single-sentence descriptions), failing to reflect complex software development scenarios. Moreover, they lack explicit requirement traceability mechanisms, making it difficult to precisely align and validate generated code against original requirements. To address these limitations, we propose TraceDev, a multi-agent framework for automated software development grounded in use cases that contain multiple functional points and complex semantics. TraceDev employs five role-specific agents, including a Requirement Refiner, Designer, Developer, Tester, and Validator. Notably, the Validator Agent constructs and maintains a heterogeneous traceability graph that links requirements, design models, and code artifacts for interacting with the preceding four agents. The traceability graph maintains consistency across various artifacts and serves as a structured context for efficient memory management, supporting reliable repository-level code generation. We evaluate TraceDev on two widely used datasets (including 125 use cases) compared with two state-of-the-art approaches. On the ETOUR dataset, TraceDev achieves a success rate of 53.63%, outperforming baseline approaches by up to 186.63%. A similar trend is observed on the SMOS dataset, where TraceDev attains a success rate of 56.82%, exceeding baseline approaches by up to 340.80%. These results demonstrate the effectiveness of TraceDev in repository-level code generation from requirements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93b5e8a5-6c7a-4754-bcf6-751d1de8a166Builds on8
- Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT ModelsJinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang et al.ICSE 2021 · 124 citations
- Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language ModelsJunhong Lin, Xinyue Zeng, Jie Zhu, Song Wang et al.ICLR 2026 · 30 citations
- RPG: A Repository Planning Graph for Unified and Scalable Codebase GenerationJane Luo, Xin Zhang, Steven Liu, Jie Wu et al.ICLR 2026 · 18 citations
- LiSSA: Toward Generic Traceability Link Recovery Through Retrieval- Augmented GenerationDominik Fuchß, Tobias Hey, Jan Keim, Haoyu Liu et al.ICSE 2025 · 8 citations
- CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklistsYukyung Lee, JoongHoon Kim, Jaehee Kim, Hyowon Cho et al.EMNLP 2025 · 2 citations
Related papers
- Towards Iterative End-to-End Software Development: A Feature-Driven Multi-agent FrameworkJunwei Liu, Chen Xu, Chong Wang, Tong Bai et al.ISSTA 2026 · 1 citation
- CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding ChallengesKechi Zhang, Jia Li, Ge Li, Xianjie Shi et al.ACL 2024
- Compiling Large Multi-modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven PerspectiveWeiyu Kong, Yun Lin, Xiwen Teoh, Duc-Minh Nguyen et al.ISSTA 2026 · 1 citation
- RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development PracticesJia Li, Hongyi Deng, Yiran Zhang, Kechi Zhang et al.FSE 2026
- PlayCoder: Making LLM-Generated GUI Code PlayableZhiyuan Peng, Wei Tao, Xin Yin, Chenhao Ying et al.FSE 2026
