Balancing Latency and Accuracy of Code Completion via Local-Cloud Model Cascading
Hanzhen Lu, Lishui Fan, Jiachi Chen, Qiuyuan Chen, Zhao Wei, Zhongxin Liu
摘要
Line-level code completion aims to complete the current line in real-time as developers type. Low latency is crucial to maintaining a seamless and uninterrupted coding experience, enabling developers to remain in a productive flow. However, existing approaches face a fundamental trade-off: large language models (LLMs) provide high-quality suggestions but require expensive computational resources to ensure acceptable inference latency. In contrast, static-analysis-based methods and small language models respond quickly but often generate suboptimal completions. To fill this gap, our idea is to rely on the small model by default and only escalate the large model when necessary to achieve latency-accuracy trade-offs. Based on this idea, we propose MCCom (Model-Cascading-based code Completion), a framework that cascades a local small model with a high-performance cloud large model for code completion. Realizing effective model cascading requires answering two non-trivial questions, i.e., when to invoke the large model and how to enable effective collaboration between small and large models. For the first question, we leverage a valuable but easily overlooked signal, i.e., user actions, during code completion to accurately identify failed completions. This deferral decision allows us to invoke the large model only when necessary, reducing both latency and cloud-side computation costs. To enable effective collaboration, MCCom employs a two-stage speculative decoding strategy and an iterative retrieval mechanism that collectively accelerate and improve the quality of completions. Due to the lack of high-quality small models for code completion, we also train a lightweight model with only 121M parameters to implement MCCom. The small model achieves an average of 73.8% of the performance of the state-of-the-art 7B model. We evaluate MCCom on the RepoEval benchmark and a new benchmark, StmtEval, collected from real-world projects. Experimental results show that our approach not only reduces inference latency by up to 47.9% and cuts down LLM usage by an average of 46.3%, but also improves the exact match rate of the large model by an average of 8.9%. CCS Concepts: • Software and its engineering → Automatic programming.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and GenerationFengji Zhang, Bei Chen, Yue Zhang, Jacky Keung 等EMNLP 2023 · 被引用 110 次
- Cascade Speculative Drafting for Even Faster LLM InferenceZiyi Chen, Xiaocong Yang, Jiacheng Lin, Chenkai Sun 等NeurIPS 2024 · 被引用 107 次
- Repoformer: Selective Retrieval for Repository-Level Code CompletionDi Wu, Wasi Uddin Ahmad, Dejiao Zhang, Murali Krishna Ramanathan 等ICML 2024 · 被引用 78 次
- RLCoder: Reinforcement Learning for Repository-Level Code CompletionYanlin Wang, Yanli Wang, Daya Guo, Jiachi Chen 等ICSE 2025 · 被引用 13 次
相关 Paper
- When Neural Code Completion Models Size up the Situation: Attaining Cheaper and Faster Completion through Dynamic Model InferenceZhensu Sun, Xiaoning Du, Fu Song, Shangwen Wang 等ICSE 2024 · 被引用 13 次
- Cascadia: An Efficient Cascade Serving System for Large Language ModelsYouhe Jiang, Fangcheng Fu, Wanru Zhao, Stephan Rabanser 等ICLR 2026 · 被引用 7 次
- Cascaded Code Editing: Large-Small Model Collaboration for Effective and Efficient Code EditingChaozheng Wang, Zezhou Yang, Shuzheng Gao, Cuiyun Gao 等FSE 2026
- AI-Assisted Code Authoring at Scale: Fine-Tuning, Deploying, and Mixed Methods EvaluationVijayaraghavan Murali, Chandra Shekhar Maddila, Imad Ahmad, Michael Bolin 等FSE 2024 · 被引用 17 次
- CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMsZhiyuan Ning, Jiawei Shao, Ruge Xu, Xinfei Guo 等NeurIPS 2025 · 被引用 5 次
