When Neural Code Completion Models Size up the Situation: Attaining Cheaper and Faster Completion through Dynamic Model Inference
Zhensu Sun, Xiaoning Du, Fu Song, Shangwen Wang, Li Li
Abstract
Leveraging recent advancements in large language models, modern neural code completion models have demonstrated the capability to generate highly accurate code suggestions. However, their massive size poses challenges in terms of computational costs and environmental impact, hindering their widespread adoption in practical scenarios. Dynamic inference emerges as a promising solution, as it allocates minimal computation during inference while maintaining the model's performance. In this research, we explore dynamic inference within the context of code completion. Initially, we conducted an empirical investigation on GPT-2, focusing on the inference capabilities of intermediate layers for code completion. We found that 54.4% of tokens can be accurately generated using just the first layer, signifying significant computational savings potential. Moreover, despite using all layers, the model still fails to predict 14.5% of tokens correctly, and the subsequent completions continued from them are rarely considered helpful, with only a 4.2% Acceptance Rate. These findings motivate our exploration of dynamic inference in code completion and inspire us to enhance it with a decision-making mechanism that stops the generation of incorrect code. We thus propose a novel dynamic inference method specifically tailored for code completion models. This method aims not only to produce correct predictions with largely reduced computation but also to prevent incorrect predictions proactively. Our extensive evaluation shows that it can averagely skip 1.7 layers out of 16 layers in the models, leading to an 11.2% speedup with only a marginal 1.1% reduction in ROUGE-L.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3d23aae-d723-450c-8773-b01d989aca97Cited by top-tier papers6
- When to Stop? Towards Efficient Code Generation in LLMs with Excess Token PreventionLianghong Guo, Yanlin Wang, Ensheng Shi, Wanjun Zhong et al.ISSTA 2024 · 16 citations
- "My productivity is boosted, but ..." Demystifying Users' Perception on AI Coding AssistantsYunbo Lyu, Zhou Yang, Jieke Shi, Jianming Chang et al.ASE 2025 · 8 citations
- API-Guided Dataset Synthesis to Finetune Large Code ModelsZongjie Li, Daoyuan Wu, Shuai Wang, Zhendong SuOOPSLA 2025 · 6 citations
- Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMsZongjie Li, Daoyuan Wu, Shuai Wang, Zhendong SuCCS 2025 · 1 citation
- Can LLM Aid in Solving Constraints with Inductive Definitions?Weizhi Feng, Shidong Shen, Jiaxiang Liu, Taolue Chen et al.FM 2026
Builds on7
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 796 citations
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang et al.NeurIPS 2020 · 401 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
- On the Importance of Building High-quality Training Datasets for Neural Code SearchZhensu Sun, Li Li, Yan Liu, Xiaoning Du et al.ICSE 2022 · 67 citations
Related papers
- D-LLM: A Token Adaptive Computing Resource Allocation Strategy for Large Language ModelsYikun Jiang, Huanyu Wang, Lei Xie, Hanbin Zhao et al.NeurIPS 2024 · 39 citations
- Balancing Latency and Accuracy of Code Completion via Local-Cloud Model CascadingHanzhen Lu, Lishui Fan, Jiachi Chen, Qiuyuan Chen et al.FSE 2026 · 1 citation
- Dynamic Context Pruning for Efficient and Interpretable Autoregressive TransformersSotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci et al.NeurIPS 2023 · 95 citations
- Skip a Layer or Loop It? Learning Program-of-Layers in LLMsZiyue Li, Yang Li, Tianyi ZhouICML 2026 · 4 citations
- EfficientEdit: Accelerating Code Editing via Edit-Oriented Speculative DecodingPeiding Wang, Li Zhang, Fang Liu, Yinghao Zhu et al.ASE 2025 · 3 citations
