Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs
Lei Zhang, Yunshui Li, Jiaming Li, Xiaobo Xia, Jiaxi Yang, Run Luo, Minzheng Wang, Longze Chen, Junhao Liu, Qiang Qu, Min Yang
摘要
Some of the latest released Code Large Language Models (Code LLMs) have been trained on repository-level code data, enabling them to perceive repository structures and utilize cross-file code information. This capability allows us to directly concatenate the content of repository code files in prompts to achieve repository-level code completion. However, in real development scenarios, directly concatenating all code repository files in a prompt can easily exceed the context window of Code LLMs, leading to a significant decline in completion performance. Additionally, overly long prompts can increase completion latency, negatively impacting the user experience. In this study, we conducted extensive experiments, including completion error analysis, topology dependency analysis, and cross-file content analysis, to investigate the factors affecting repository-level code completion. Based on the conclusions drawn from these preliminary experiments, we proposed a strategy called Hierarchical Context Pruning (HCP) to construct high-quality completion prompts. We applied the HCP to six Code LLMs and evaluated them on the CrossCodeEval dataset. The experimental results showed that, compared to previous methods, the prompts constructed using our HCP strategy achieved higher completion accuracy on five out of six Code LLMs. Additionally, the HCP managed to keep the prompt length around 8k tokens (whereas the full repository code is approximately 50k tokens), significantly improving completion throughput. Our code and data will be publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Continual Multimodal Contrastive LearningXiaohao Liu, Xiaobo Xia, See-Kiong Ng, Tat-Seng ChuaNeurIPS 2025 · 被引用 25 次
- L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language ModelsXiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang 等NeurIPS 2025 · 被引用 16 次
- OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech SynthesisRun Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu 等NeurIPS 2025 · 被引用 5 次
- Chiseling Out Efficiency: Structured Skeleton Supervision for Efficient Code GenerationYu Yu, Zhihong Sun, Jia Li, Yao Wan 等FSE 2026
- In Line with Context: Repository-Level Code Generation via Context InliningChao Hu, Wenhao Zeng, Yuling Shi, Beijun Shen 等FSE 2026
它引用的顶会 Paper7
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Repository-Level Prompt Generation for Large Language Models of CodeDisha Shrivastava, Hugo Larochelle, Daniel TarlowICML 2023 · 被引用 184 次
- Lemur: Harmonizing Natural Language and Code for Language AgentsYiheng Xu, Hongjin Su, Chen Xing, Boyu Mi 等ICLR 2024 · 被引用 92 次
- EHRAgent: Code Empowers Large Language Models for Few-shot Complex Tabular Reasoning on Electronic Health RecordsWenqi Shi, Ran Xu, Yuchen Zhuang, Yue Yu 等EMNLP 2024 · 被引用 33 次
相关 Paper
- Aligning LLMs to Fully Utilize the Cross-file Context in Repository-level Code CompletionJia Li, Hao Zhu, Huanyu Liu, Xianjie Shi 等ASE 2025 · 被引用 2 次
- M2RC-EVAL: Massively Multilingual Repository-level Code Completion EvaluationJiaheng Liu, Ken Deng, Congnan Liu, Jian Yang 等ACL 2025 · 被引用 19 次
- AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code CompletionTianyue Jiang, Yanlin Wang, Yanli Wang, Daya Guo 等ASE 2025 · 被引用 2 次
- CodeRAG: Finding Relevant and Necessary Knowledge for Retrieval-Augmented Repository-Level Code CompletionSheng Zhang, Yifan Ding, Shuquan Lian, Shun Song 等EMNLP 2025 · 被引用 3 次
- RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsTianyang Liu, Canwen Xu, Julian J. McAuleyICLR 2024 · 被引用 338 次
