What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and Beyond
Wenchao Gu, Juntao Chen, Yanlin Wang, Tianyue Jiang, Xingzhe Li, Mingwei Liu, Xilin Liu, Yuchi Ma, Zibin Zheng
摘要
Repository-level code generation remains challenging due to complex code dependencies and the limitations of large language models (LLMs) in processing long contexts. While retrieval-augmented generation (RAG) frameworks are widely adopted, the effectiveness of different retrieved information sources-contextual code, APIs, and similar snippets-has not been rigorously analyzed. Through an empirical study on two benchmarks, we demonstrate that in-context code and potential API information significantly enhance LLM performance, whereas retrieved similar code often introduces noise, degrading results by up to 15%. Based on the preliminary results, we propose AllianceCoder, a novel context-integrated method that employs chain-of-thought prompting to decompose user queries into implementation steps and retrieves APIs via semantic description matching. Through extensive experiments on CoderEval and RepoExec, AllianceCoder achieves state-of-the-art performance, improving Pass@1 by up to 20% over existing approaches. This study provides an experimental framework to further exploring what to retrieve in RAG-based code generation, with our replication package available at https://anonymous.4open.science/r/AllianceCoder to facilitate future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- RESCUE: Retrieval Augmented Secure Code GenerationJiahao Shi, Tianyi ZhangICLR 2026 · 被引用 15 次
- DrainCode: Stealthy Energy Consumption Attacks on Retrieval-Augmented Code Generation via Context PoisoningYanli Wang, Jiadong Wu, Tianyue Jiang, Mingwei Liu 等ASE 2025 · 被引用 3 次
- AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code CompletionTianyue Jiang, Yanlin Wang, Yanli Wang, Daya Guo 等ASE 2025 · 被引用 2 次
- RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language ModelsYanlin Wang, Suiquan Wang, Yanli Wang, Bowen Zhang 等FSE 2026
- Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and BeyondMinh Le-Anh, Huyen Nguyen, Khanh An Tran, Nam Le Hai 等FSE 2026
它引用的顶会 Paper25
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Making Retrieval-Augmented Language Models Robust to Irrelevant ContextOri Yoran, Tomer Wolfson, Ori Ram, Jonathan BerantICLR 2024 · 被引用 361 次
- RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsTianyang Liu, Canwen Xu, Julian J. McAuleyICLR 2024 · 被引用 338 次
相关 Paper
- RepoScope: Leveraging Call Chain-Aware Multi-View Context for Repository-Level Code GenerationYang Liu, Li Zhang, Fang Liu, Zhuohang Wang 等ICSE 2026
- In Line with Context: Repository-Level Code Generation via Context InliningChao Hu, Wenhao Zeng, Yuling Shi, Beijun Shen 等FSE 2026
- GraphCoder: Enhancing Repository-Level Code Completion via Coarse-to-fine Retrieval Based on Code Context GraphWei Liu, Ailun Yu, Daoguang Zan, Bo Shen 等ASE 2024 · 被引用 11 次
- Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'Shanchao Liang, Nan Jiang, Yiran Hu, Lin TanACL 2025 · 被引用 9 次
- SRACG: A Code Generation Framework with Selective Retrieval AugmentationMengzhen Wang, Shukai Ma, Songwen Gong, Jiexin Wang 等AAAI 2026
