RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
Jane Luo, Xin Zhang, Steven Liu, Jie Wu, Jianfeng Liu, Yiming Huang, Yangyu Huang, Chengyu Yin, Ying Xin, Yuefeng Zhan, Hao Sun, Qi Chen
Abstract
Large language models excel at generating individual functions or single files of code, yet generating complete repositories from scratch remains a fundamental challenge. This capability is key to building coherent software systems from highlevel specifications and realizing the full potential of automated code generation. The process requires planning at two levels: deciding what features and modules to build (proposal stage) and defining their implementation details (implementation stage). Current approaches rely on natural language planning, which often produces unclear specifications, misaligned components, and brittle designs due to its inherent ambiguity and lack of structure. To address these limitations, we introduce the Repository Planning Graph (RPG), a structured representation that encodes capabilities, file structures, data flows, and functions in a unified graph. By replacing free-form natural language with an explicit blueprint, RPG enables consistent long-horizon planning for repository generation. Building on RPG, we develop ZeroRepo, a graph-driven framework that operates in three stages: proposal-level planning, implementation-level construction, and graph-guided code generation with test validation To evaluate, we construct RepoCraft, a benchmark of six real-world projects with 1,052 tasks. On RepoCraft, ZeroRepo produces nearly 36K Code Lines and 445K Code Tokens, on average 3.9× larger than the strongest baseline (Claude Code), and 68× larger than other baselines. It achieves 81.5% coverage and 69.7% test accuracy, improving over Claude Code by 27.3 and 35.8 points. Further analysis shows that RPG models complex dependencies, enables more sophisticated planning through near-linear scaling, and improves agent understanding of repositories, thus accelerating localization. Our data and code are available at https://github.com/microsoft/RPG-ZeroRepo . INTRODUCTION Recent large language models (LLMs) have shown strong performance on function-level and file-level code generation, reliably producing functions and files from natural language descriptions (Zhu et al., 2024; Wang et al., 2025; Liu et al., 2025; Zeng et al., 2025) . However, scaling this capability from functions and files to generate large-scale software repositories from scratch remains a fundamental challenge. The core difficulty is bridging the gap between high-level user intent and the repository's intricate network of files, classes, and dependenciesTao et al. (2025); Li (2025). Successfully navigating this gap necessitates a process of progressive planning, which naturally decomposes into two complementary phases: proposal-level planning, which determines what to build by defining the functional scope and key capabilities, and implementation-level planning, which determines how to build it by specifying the file structure, interfaces, dependencies, and data flows.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27696198-3f4e-4de0-b88a-6fd48c0cc663Cited by top-tier papers5
- TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test GenerationSteven Liu, Jane Luo, Xin Zhang, Aofan Liu et al.ICML 2026 · 4 citations
- Towards Iterative End-to-End Software Development: A Feature-Driven Multi-agent FrameworkJunwei Liu, Chen Xu, Chong Wang, Tong Bai et al.ISSTA 2026 · 1 citation
- NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into CodeSeemandhar Jain, Keshav Gupta, Kunal Gupta, Manmohan ChandrakerCVPR 2026 · 1 citation
- Closing the Loop: Universal Repository Representation with RPG-EncoderJane Luo, Chengyu Yin, Xin Zhang, Qingtao Li et al.ICML 2026
- TraceDev: A Traceability-Driven Multi-agent Framework for Requirement-to-Code DevelopmentMingyu Chen, Yakun Zhang, Zihao Xie, Yixing Luo et al.ISSTA 2026
Builds on12
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- AdaPlanner: Adaptive Planning from Feedback with Language ModelsHaotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai et al.NeurIPS 2023 · 257 citations
- Paper2Code: Automating Code Generation from Scientific Papers in Machine LearningMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangICLR 2026 · 86 citations
- rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified DatasetYifei Liu, Li Lyna Zhang, Yi Zhu, Bingcheng Dong et al.NeurIPS 2025 · 50 citations
- Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering TasksHongyuan Tao, Ying Zhang, Zhenhao Tang, Hongen Peng et al.NeurIPS 2025 · 41 citations
Related papers
- Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'Shanchao Liang, Nan Jiang, Yiran Hu, Lin TanACL 2025 · 9 citations
- RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development PracticesJia Li, Hongyi Deng, Yiran Zhang, Kechi Zhang et al.FSE 2026
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding AgentsJingzhe Ding, Shengda Long, Changxin Pu, Ge Zhang et al.ICML 2026 · 37 citations
- RepoScope: Leveraging Call Chain-Aware Multi-View Context for Repository-Level Code GenerationYang Liu, Li Zhang, Fang Liu, Zhuohang Wang et al.ICSE 2026
- Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and BeyondMinh Le-Anh, Huyen Nguyen, Khanh An Tran, Nam Le Hai et al.FSE 2026
