Balancing Latency and Accuracy of Code Completion via Local-Cloud Model Cascading
Hanzhen Lu, Lishui Fan, Jiachi Chen, Qiuyuan Chen, Zhao Wei, Zhongxin Liu
Abstract
Line-level code completion aims to complete the current line in real-time as developers type. Low latency is crucial to maintaining a seamless and uninterrupted coding experience, enabling developers to remain in a productive flow. However, existing approaches face a fundamental trade-off: large language models (LLMs) provide high-quality suggestions but require expensive computational resources to ensure acceptable inference latency. In contrast, static-analysis-based methods and small language models respond quickly but often generate suboptimal completions. To fill this gap, our idea is to rely on the small model by default and only escalate the large model when necessary to achieve latency-accuracy trade-offs. Based on this idea, we propose MCCom (Model-Cascading-based code Completion), a framework that cascades a local small model with a high-performance cloud large model for code completion. Realizing effective model cascading requires answering two non-trivial questions, i.e., when to invoke the large model and how to enable effective collaboration between small and large models. For the first question, we leverage a valuable but easily overlooked signal, i.e., user actions, during code completion to accurately identify failed completions. This deferral decision allows us to invoke the large model only when necessary, reducing both latency and cloud-side computation costs. To enable effective collaboration, MCCom employs a two-stage speculative decoding strategy and an iterative retrieval mechanism that collectively accelerate and improve the quality of completions. Due to the lack of high-quality small models for code completion, we also train a lightweight model with only 121M parameters to implement MCCom. The small model achieves an average of 73.8% of the performance of the state-of-the-art 7B model. We evaluate MCCom on the RepoEval benchmark and a new benchmark, StmtEval, collected from real-world projects. Experimental results show that our approach not only reduces inference latency by up to 47.9% and cuts down LLM usage by an average of 46.3%, but also improves the exact match rate of the large model by an average of 8.9%. CCS Concepts: • Software and its engineering → Automatic programming.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e77232d2-8b89-4087-b55e-a63aa02810ccBuilds on9
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and GenerationFengji Zhang, Bei Chen, Yue Zhang, Jacky Keung et al.EMNLP 2023 · 110 citations
- Cascade Speculative Drafting for Even Faster LLM InferenceZiyi Chen, Xiaocong Yang, Jiacheng Lin, Chenkai Sun et al.NeurIPS 2024 · 107 citations
- Repoformer: Selective Retrieval for Repository-Level Code CompletionDi Wu, Wasi Uddin Ahmad, Dejiao Zhang, Murali Krishna Ramanathan et al.ICML 2024 · 78 citations
- RLCoder: Reinforcement Learning for Repository-Level Code CompletionYanlin Wang, Yanli Wang, Daya Guo, Jiachi Chen et al.ICSE 2025 · 13 citations
Related papers
- When Neural Code Completion Models Size up the Situation: Attaining Cheaper and Faster Completion through Dynamic Model InferenceZhensu Sun, Xiaoning Du, Fu Song, Shangwen Wang et al.ICSE 2024 · 13 citations
- Cascadia: An Efficient Cascade Serving System for Large Language ModelsYouhe Jiang, Fangcheng Fu, Wanru Zhao, Stephan Rabanser et al.ICLR 2026 · 7 citations
- Cascaded Code Editing: Large-Small Model Collaboration for Effective and Efficient Code EditingChaozheng Wang, Zezhou Yang, Shuzheng Gao, Cuiyun Gao et al.FSE 2026
- AI-Assisted Code Authoring at Scale: Fine-Tuning, Deploying, and Mixed Methods EvaluationVijayaraghavan Murali, Chandra Shekhar Maddila, Imad Ahmad, Michael Bolin et al.FSE 2024 · 17 citations
- CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMsZhiyuan Ning, Jiawei Shao, Ruge Xu, Xinfei Guo et al.NeurIPS 2025 · 5 citations
