MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache Optimization
Borui Li, Yitao Wang, Haoran Ma, Ligeng Chen, Jun Xiao, Shuai Wang
Abstract
Deploying large language models (LLMs) with low-rank adaptation (LoRA) on mobile devices is promising due to their capability to complete diverse domain-specific tasks while ensuring privacy and accessibility. In this paper, we introduce MobiLoRA to accelerate LoRA-based LLM inference on mobile devices. MobiLoRA focuses on optimizing the key-value (KV) caches due to the limited computing and memory resources of mobile devices. The key insight of MobiLoRA lies in the utilization of two contexts for on-device LoRA serving: semantic-level contexts, such as prompts with shared prefixes, and system-level contexts, such as the application status (e.g., foreground or killed) of LLM requests. Specifically, for semantic-level contexts, Mo-biLoRA proposes similarity-aware delta encoding, which leverages token-wise similarity in KV caches across LoRA adapters for efficient storage and reuse. Furthermore, Mo-biLoRA advocates context-aware KV cache management to optimize cache eviction considering the system-level contexts. We implement MobiLoRA and compare it with state-of-the-art LLM serving frameworks using real-world mobile device traces. Results show that Mo-biLoRA accelerates LoRA-based LLM inference by 18.1% 80.5% on mobile devices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 304d1a6a-3039-4356-9dd8-b17bc44a6e1fCited by top-tier papers2
- LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM AgentsHyesung Jeon, Hyeongju Ha, jae-joon kimICML 2026 · 5 citations
- When KV Cache Reuse Fails in Multi-Agent Systems: Cross-Candidate Interaction is Crucial for LLM JudgesSichu Liang, Zhenglin Wang, Jiajia Chu, Pengfei Xia et al.ACL 2026
Builds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Parrot: Efficient Serving of LLM-based Applications with Semantic VariableChaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang et al.OSDI 2024 · 112 citations
Related papers
- SemCache: Semantic-Aware Cache Sharing for Efficient Multi-User LoRA-Adapted LLM Inference at the EdgeTao Ren, Yiming Yao, Zheyuan Hu, Jianwei NiuINFOCOM 2026 · 1 citation
- ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM ServingJiuchen Shi, Hang Zhang, Yixiao Wang, Quan Chen et al.HPCA 2026
- dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM ServingBingyang Wu, Ruidong Zhu, Zili Zhang, Peng Sun et al.OSDI 2024 · 79 citations
- Empower Vision Applications with LoRA LMMLiang Mi, Weijun Wang, Wenming Tu, Qingfeng He et al.EuroSys 2025 · 2 citations
- Activated LoRA: Fine-tuned LLMs for IntrinsicsKristjan Greenewald, Luis A. Lastras, Thomas Parnell, Vraj Shah et al.NeurIPS 2025 · 12 citations
