LiLo: Harnessing the on-Chip Accelerators in Intel CPUs for Compressed LLM Inference Acceleration
Hyungyo Kim, Qirong Xia, Jinghan Huang, Nachuan Wang, Younjoo Lee, Jung Ho Ahn, Wajdi K. Feghali, Ren Wang, Nam Sung Kim
Abstract
The ever-growing sizes of large language models (LLMs) introduce significant infrastructure challenges due to their immense memory capacity demands. While the de facto approach has been to deploy multiple high-end GPUs, each with a limited memory capacity, the prohibitive cost of such systems has become a major barrier to the widespread deployment of frontier LLMs. As a result, CPU-based inference has become an appealing and cost-efficient alternative, since a CPU can offer an order of magnitude larger memory capacity at a fraction of the cost while providing competitive throughput for matrixvector multiplication with the latest Advanced Matrix Extensions (AMX). It not only broadens accessibility for users without multiGPU setups but also enables hyperscalers to leverage underutilized CPU servers to accommodate temporarily surging inference demand. Nevertheless, even CPU's large memory capacity has become insufficient to serve LLMs with hundreds of billions of parameters. Under the memory capacity constraint, we may offload parameters to storage devices and fetch them on demand, but doing so significantly degrades inference performance due to the high latency and low bandwidth of storage devices. To address this challenge, we propose LILO, an LLM inference framework that leverages In-memory Analytics Accelerator (IAA) in the latest Intel CPUs, to accelerate inference under memory capacity constraints. By storing model parameters in a compressed format and decompressing them on demand using IAA, LILO enables significantly reduced storage access during inference under memory capacity constraints while preserving the model accuracy and behavior. LILO orchestrates the concurrent execution of on-chip accelerators, i.e., IAA, Advanced Vector Extensions (AVX), and AMX, to facilitate high-throughput decompression alongside inference computation. Furthermore, LILO implements selective compression, a Mixture-of-Expert (MoE)-aware optimization that reduces the decompression overhead by up to 1.9×. We demonstrate that LILO reduces inference latency by up to 4.9× and 4.3× for Llama3-405B and DeepSeekR1, respectively, under memory capacity constraints compared to the baseline inference solely relying on storage-offloading without compression.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 10c657cc-2b42-44a9-9af2-66f3735882c7Cited by top-tier papers1
Ask how each one uses itRelated papers
- DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline ModelGerasimos Gerogiannis, Stijn Eyerman, Evangelos Georganas, Wim Heirman et al.MICRO 2025 · 5 citations
- LIA: A Single-GPU LLM Inference Acceleration with Cooperative AMX-Enabled CPU-GPU Computation and CXL OffloadingHyungyo Kim, Nachuan Wang, Qirong Xia, Jinghan Huang et al.ISCA 2025 · 20 citations
- Towards Resource-Efficient Serverless LLM Inference with SLINFERChuhao Xu, Zijun Li, Quan Chen, Han Zhao et al.HPCA 2026
- DynamicInfer: Runtime-Aware Sparse Offloading for LLMs Inference on a Consumer-Grade GPUZhui Zhu, Weichen Zhang, Zhenghan Zhou, Yunhao Liu et al.ICLR 2026
- Practical Offloading for Fine-Tuning LLM on Commodity GPU via Learned Sparse ProjectorsSiyuan Chen, Zhuofeng Wang, Zelong Guan, Yudong Liu et al.AAAI 2025 · 3 citations
