Lune

FSE2026顶会

Compiling Code LLMs into Lightweight Executables

Jieke Shi, Junda He, Zhou Yang, Chengran Yang, Mykhailo V. Klymenko, Thong Hoang, Xiwei Xu, Zhenchang Xing, David Lo

2026年份

摘要

The demand for better prediction accuracy and higher execution performance in neural networks continues to grow. The emergence and success of Large Language Models (LLMs) have produced many cloud-based tools for software engineering tasks such as code suggestion. Although effective, cloud deployment raises concerns over privacy, latency, and reliance on network connectivity. Running LLMs locally on personal devices such as laptops would address these issues, because it enables offline use and reduces response time. However, local deployment is challenging, since commodity devices lack high-performance accelerators such as GPUs and are constrained by limited memory and compute capacity, which makes it hard to execute large models efficiently.

We present Ditto, a framework that optimizes both the model size of Code LLMs and the inference programs that execute them, with a focus on statically-typed languages such as C. Our approach integrates two components. The first is a quantization technique inspired by product quantization, which groups model parameters into per-block codebooks via K-Means clustering and stores each weight as a bit-packed low-bitwidth index. The quantizer further supports a mixed-precision mode that keeps a small number of sensitivity-critical tensors in float32. The second component is a compilation pass integrated into LLVM that automatically detects and replaces unoptimized General Matrix-Vector Multiplication (GEMV) operations, which are the most computationally intensive part of code models, with calls into Basic Linear Algebra Subprograms (BLAS) libraries that are highly optimized for the target hardware. The output of Ditto is a compiled executable that runs the selected Code LLM on commodity hardware.

We evaluate Ditto on three popular Code LLMs, namely Code Llama, MagicCoder, and OpenCodeInterpreter, achieving up to 10.5× faster inference, 6.4× lower memory usage, and 10.5× lower energy consumption compared with their original inference pipelines, while preserving accuracy close to the full-precision models, with an average loss of only 0.27% in pass@1. Ditto also outperforms the state-of-the-art int8 quantization baseline, achieving up to 6.96% higher pass@1 accuracy, 2.2× speedup, and 1.6× memory reduction, which demonstrates the effectiveness of our approach.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext de43b77e-7ae5-480e-97ab-e710cb6805e5

它引用的顶会 Paper21

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖