LIA: A Single-GPU LLM Inference Acceleration with Cooperative AMX-Enabled CPU-GPU Computation and CXL Offloading
Hyungyo Kim, Nachuan Wang, Qirong Xia, Jinghan Huang, Amir Yazdanbakhsh, Nam Sung Kim
Abstract
The limited memory capacity of single GPUs constrains large language model (LLM) inference, necessitating cost-prohibitive multi-GPU deployments or frequent performance-limiting CPU-GPU transfers over slow PCIe.In this work, we first benchmark recent Intel CPUs with Advanced Matrix Extensions (AMX), including 4th generation (Sapphire Rapids) and 6th generation (Granite Rapids) Xeon Scalable Processors, demonstrating matrix multiplication throughput of 20 TFLOPS and 40 TFLOPS, respectivelycomparable to some recent GPUs.These findings unlock more extensive computation offloading to CPUs, reducing CPU-GPU transfers and alleviating throughput bottlenecks compared to priorgeneration CPUs.Building on these insights, we design LIA, a single-GPU LLM inference acceleration framework leveraging cooperative AMX-enabled CPU-GPU computation and CXL offloading.LIA systematically offloads computation to CPUs, optimizing both latency and throughput.The framework also introduces a memoryoffloading policy that seamlessly integrates affordable CXL memory with DDR memory to enhance performance in throughput-driven tasks.On Saphhire Rapids (Granite Rapids) systems with a single H100 GPU, LIA achieves up to 5.1× (19×) lower latency and 3.7× (5.1×) higher throughput compared to the latest single-GPU offloading framework.Furthermore, LIA deploying CXL offloading yields an additional 1.5× throughput improvement over LIA using only DDR memory with a 1.8× increase in maximum batch size (900→1.6K).
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 68fc7796-936d-4fdd-a626-6c136e68d8ddCited by top-tier papers8
- SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert SubstitutionGuoying Zhu, Meng Li, Haipeng Dai, Xuechen Liu et al.ISCA 2026 · 4 citations
- AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM ServingXinkai Wang, Chao Li, Yiming Zhuansun, Jinyang Guo et al.HPCA 2026 · 2 citations
- FastTTS: Accelerating Test-Time Scaling for Edge LLM ReasoningHao Mark Chen, Zhiwen Mo, Guanxi Lu, Shuang Liang et al.ASPLOS 2026 · 1 citation
- Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention PiggybackingZizhao Mo, Junlin Chen, Huanle Xu, ChengZhong XuSIGMOD 2026 · 1 citation
- Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System SimulationYanjing Wang, Lizhou Wu, Sunfeng Gao, Yibo Tang et al.HPCA 2026 · 1 citation
Related papers
- LiLo: Harnessing the on-Chip Accelerators in Intel CPUs for Compressed LLM Inference AccelerationHyungyo Kim, Qirong Xia, Jinghan Huang, Nachuan Wang et al.HPCA 2026 · 1 citation
- DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline ModelGerasimos Gerogiannis, Stijn Eyerman, Evangelos Georganas, Wim Heirman et al.MICRO 2025 · 5 citations
- Practical Offloading for Fine-Tuning LLM on Commodity GPU via Learned Sparse ProjectorsSiyuan Chen, Zhuofeng Wang, Zelong Guan, Yudong Liu et al.AAAI 2025 · 3 citations
- Understand and Accelerate Memory Processing Pipeline for Large Language Model InferenceZifan He, Rui Ma, Yizhou Sun, Jason CongICML 2026
- High-Throughput Non-uniformly Quantized 3-bit LLM InferenceYuAng Chen, Wenqi Zeng, Jeffrey Xu YuPPoPP 2026
