Accelerating Model Loading in LLM Inference by Programmable Page Cache
Yubo Liu, Hongbo Li, Xiaojia Huang, Yongfeng Wang, Hanjun Guo, Hui Chen, Yuxin Ren, Ning Jia
摘要
This paper examines the model loading bottleneck during the LLM inference startup. Existing solutions often optimize model loading performance at the expense of compatibility. However, compatibility is a crucial factor determining whether a technology can be widely applied in real-world scenarios. This work achieves both high performance and strong compatibility by optimizing the cache policy of the kernel file system. We design PPC, a programmable page cache framework that allows users to customize page cache policies in a non-intrusive, flexible, and lightweight manner. Furthermore, we design MAIO, a cache policy implemented based on PPC, to optimize model loading. MAIO introduces an I/O template-based mechanism to fully utilize SSD bandwidth, XPU affinity, and data locality to enhance the efficiency of prefetching and eviction. Our evaluation shows that MAIO reduces the model loading latency by up to 79% compared to existing optimizations. In a real-world application, MAIO achieves up to 36% improvement in inference throughput over other tested solutions in the elastic deployment scenario.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- INFless: a native serverless system for low-latency, high-throughput inferenceYanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang 等ASPLOS 2022 · 被引用 145 次
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete 等OSDI 2024 · 被引用 125 次
- DADI: Block-Level Image Service for Agile and Elastic Application DeploymentHuiba Li, Yifan Yuan, Rui Du, Kai Ma 等USENIX ATC 2020 · 被引用 56 次
相关 Paper
- Compute or Load KV Cache? Why Not Both?Shuowei Jin, Xueshen Liu, Qingzhao Zhang, Zhuoqing MaoICML 2025
- PASK: Cold Start Mitigation for Inference with Proactive and Selective Kernel Loading on GPUsXuanteng Huang, Jiangsu Du, Nong Xiao, XianWei ZhangDAC 2025
- LiLo: Harnessing the on-Chip Accelerators in Intel CPUs for Compressed LLM Inference AccelerationHyungyo Kim, Qirong Xia, Jinghan Huang, Nachuan Wang 等HPCA 2026 · 被引用 1 次
- PF-LLM: Large Language Model Hinted Hardware PrefetchingCeyu Xu, Xiangfeng Sun, Weihang Li, Chen Bai 等ASPLOS 2026
- REPA: Reconfigurable PIM for the Joint Acceleration of KV Cache Offloading and ProcessingYang Hong, Junlong Yang, Bo Peng, Jianguo YaoASPLOS 2026
