Accelerating Model Loading in LLM Inference by Programmable Page Cache
Yubo Liu, Hongbo Li, Xiaojia Huang, Yongfeng Wang, Hanjun Guo, Hui Chen, Yuxin Ren, Ning Jia
Abstract
This paper examines the model loading bottleneck during the LLM inference startup. Existing solutions often optimize model loading performance at the expense of compatibility. However, compatibility is a crucial factor determining whether a technology can be widely applied in real-world scenarios. This work achieves both high performance and strong compatibility by optimizing the cache policy of the kernel file system. We design PPC, a programmable page cache framework that allows users to customize page cache policies in a non-intrusive, flexible, and lightweight manner. Furthermore, we design MAIO, a cache policy implemented based on PPC, to optimize model loading. MAIO introduces an I/O template-based mechanism to fully utilize SSD bandwidth, XPU affinity, and data locality to enhance the efficiency of prefetching and eviction. Our evaluation shows that MAIO reduces the model loading latency by up to 79% compared to existing optimizations. In a real-world application, MAIO achieves up to 36% improvement in inference throughput over other tested solutions in the elastic deployment scenario.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11c01743-7a87-4213-b686-0945fc089220Builds on15
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- INFless: a native serverless system for low-latency, high-throughput inferenceYanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang et al.ASPLOS 2022 · 145 citations
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete et al.OSDI 2024 · 125 citations
- DADI: Block-Level Image Service for Agile and Elastic Application DeploymentHuiba Li, Yifan Yuan, Rui Du, Kai Ma et al.USENIX ATC 2020 · 56 citations
Related papers
- Compute or Load KV Cache? Why Not Both?Shuowei Jin, Xueshen Liu, Qingzhao Zhang, Zhuoqing MaoICML 2025
- PASK: Cold Start Mitigation for Inference with Proactive and Selective Kernel Loading on GPUsXuanteng Huang, Jiangsu Du, Nong Xiao, XianWei ZhangDAC 2025
- LiLo: Harnessing the on-Chip Accelerators in Intel CPUs for Compressed LLM Inference AccelerationHyungyo Kim, Qirong Xia, Jinghan Huang, Nachuan Wang et al.HPCA 2026 · 1 citation
- PF-LLM: Large Language Model Hinted Hardware PrefetchingCeyu Xu, Xiangfeng Sun, Weihang Li, Chen Bai et al.ASPLOS 2026
- REPA: Reconfigurable PIM for the Joint Acceleration of KV Cache Offloading and ProcessingYang Hong, Junlong Yang, Bo Peng, Jianguo YaoASPLOS 2026
