One Size Does Not Fit All: Revisiting Code Context Engineering for Repository-Level Code Generation
Yichen Li, Qiye Lin, Yun Peng, Zhihan Jiang, Jinyang Liu, Chaozheng Wang, Yintong Huo, Cuiyun Gao
摘要
Large Language Models (LLMs) have demonstrated remarkable capabilities in code generation at the function or file level. However, they achieve limited performance on repository-level code generation due to the complicated repository context where a substantial amount of files and functions exist. To address this challenge, code context engineering methods are proposed to accurately extract the relevant code context required by such tasks. These methods belong to three dominant paradigms: 1) Similarity-based In-Context Learning (ICL) , which retrieves similar code examples in the repository as demonstrations; 2) Static Analysis , which captures relevant code context based on structural dependencies in the repository; and 3) Navigation-based Paradigms , which invokes LLMs to dynamically explore repositories for context identification. Despite the prevalence of code context engineering approaches, their individual contributions are often coupled within complex agents in repository-level code generation, making it difficult to isolate and evaluate the effectiveness of each paradigm. This paper presents the first large-scale and systematic empirical study that compares three code context engineering paradigms independently. We evaluate ten representative code context engineering methods based on the three paradigms, with eight popular LLMs on this task. We also propose a new metric named Dependency Collection Rate (DCR) and efficiency metrics to enable the direct comparison of code context engineering methods, rather than observing their impacts only on the final code generation performance. Our findings reveal fundamental trade-offs: static analysis provides the most reliable balance between effectiveness and cost for function-level generation, while navigation-based approaches become increasingly advantageous as task complexity grows. However, navigation requires powerful models and incurs 10-20× higher computational costs. Based on these findings, we provide actionable implications for AI coding researchers and software developers to guide the design and deployment of context-aware coding tools.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and BeyondWenchao Gu, Juntao Chen, Yanlin Wang, Tianyue Jiang 等ICSE 2026 · 被引用 1 次
- LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and MitigationZiyao Zhang, Chong Wang, Yanlin Wang, Ensheng Shi 等ISSTA 2025 · 被引用 53 次
- HumanEvo: An Evolution-Aware Benchmark for More Realistic Evaluation of Repository-Level Code GenerationDewu Zheng, Yanlin Wang, Ensheng Shi, Ruikai Zhang 等ICSE 2025 · 被引用 2 次
- Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and BeyondMinh Le-Anh, Huyen Nguyen, Khanh An Tran, Nam Le Hai 等FSE 2026
- What Makes Good In-Context Demonstrations for Code Intelligence Tasks with LLMs?Shuzheng Gao, Xin-Cheng Wen, Cuiyun Gao, Wenxuan Wang 等ASE 2023 · 被引用 80 次
