Better Call Grep: Evaluating and Improving Grep-Like Lexical Retrieval for Repository-Level Code Completion
Baoyi Wang, Xingliang Wang, Guochang Li, Chen Zhi, Junxiao Han, Xinkui Zhao, Nan Wang, Shuiguang Deng, Jianwei Yin
摘要
Repository-level code completion remains challenging for large language models (LLMs), as it requires reasoning over cross-file dependencies while under limited context windows. To address this challenge, prior work has adopted Retrieval-Augmented Generation (RAG) frameworks based on semantic indexing or structureaware graph analysis. Although effective, these approaches introduce substantial computational overhead for index construction and maintenance, which hinders their practicality in real-world development. Motivated by common developer workflows that rely on lightweight search utilities (e.g., ripgrep) to locate relevant code, we revisit a fundamental yet underexplored question: how far can simple, index-free lexical retrieval go in supporting repository-level code completion before more complex retrieval mechanisms become necessary? To answer this question, we systematically explore the potential of lightweight, index-free, intent-aware lexical retrieval through extensive empirical analysis. We first introduce Naive GrepRAG, a baseline framework where LLMs autonomously generate ripgrep commands to localize relevant context. Our preliminary experiments show that even this basic implementation achieves performance comparable to sophisticated graph-based baselines. Further analysis reveals that its effectiveness stems from retrieving code fragments that are lexically precise and spatially closer to the completion site. However, we identify key limitations of this approach, including sensitivity to noisy matches caused by high-frequency ambiguous keywords and context fragmentation due to rigid truncation boundaries. To address these issues, we propose GrepRAG, which augments lexical retrieval with a lightweight post-processing pipeline featuring identifier-weighted re-ranking and structure-aware deduplication. Extensive evaluation on CrossCodeEval and RepoEval_Updated demonstrates that GrepRAG consistently outperforms state-of-the-art (SOTA) methods. In particular, on CrossCodeEval, GrepRAG achieves 7.04-15.58% relative improvement in code exact match (EM) over the best baseline. CCS Concepts: • Software and its engineering → Software development techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
- RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsTianyang Liu, Canwen Xu, Julian J. McAuleyICLR 2024 · 被引用 338 次
- ReACC: A Retrieval-Augmented Code Completion FrameworkShuai Lu, Nan Duan, Hojae Han, Daya Guo 等ACL 2022 · 被引用 208 次
- Repository-Level Prompt Generation for Large Language Models of CodeDisha Shrivastava, Hugo Larochelle, Daniel TarlowICML 2023 · 被引用 184 次
相关 Paper
- Repoformer: Selective Retrieval for Repository-Level Code CompletionDi Wu, Wasi Uddin Ahmad, Dejiao Zhang, Murali Krishna Ramanathan 等ICML 2024 · 被引用 78 次
- CodeRAG: Finding Relevant and Necessary Knowledge for Retrieval-Augmented Repository-Level Code CompletionSheng Zhang, Yifan Ding, Shuquan Lian, Shun Song 等EMNLP 2025 · 被引用 3 次
- AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code CompletionTianyue Jiang, Yanlin Wang, Yanli Wang, Daya Guo 等ASE 2025 · 被引用 2 次
- Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and BeyondMinh Le-Anh, Huyen Nguyen, Khanh An Tran, Nam Le Hai 等FSE 2026
- GraphCoder: Enhancing Repository-Level Code Completion via Coarse-to-fine Retrieval Based on Code Context GraphWei Liu, Ailun Yu, Daoguang Zan, Bo Shen 等ASE 2024 · 被引用 11 次
