SemCache: Semantic-Aware Cache Sharing for Efficient Multi-User LoRA-Adapted LLM Inference at the Edge
Tao Ren, Yiming Yao, Zheyuan Hu, Jianwei Niu
Abstract
Large language models (LLMs) enable personalized applications through parameter-efficient techniques like Low-Rank Adaptation (LoRA). Deploying LoRA-adapted LLMs faces a critical dilemma: resource-constrained user devices (UDs) are hard to host full-scale LLMs, while centralized edge deployments of all LoRA adapters incur huge storage and privacy risks. We propose EdgeLoRA, a collaborative framework where UDs retain private LoRA adapters locally while offloading base model computations to edge servers (ESs), enhancing privacy and reducing storage overhead. However, EdgeLoRA introduces latency inefficiency from multi-round ES-UD interactions and memory overhead from redundant per-user key-value (KV) caching. To address this, we design SemCache, a semantic-aware caching mechanism that exploits hierarchical redundancies. SemCache first clusters user queries into coarse-grained intents (e.g., health, economic) via lightweight textual encoders, then identifies reusable fine-grained subsequences (e.g., "feel pain in") within clusters. A hybrid-indexed global cache (cluster index + token position) enables cross-user KV reuse, while an adaptive policy optimizes cache utility via frequency, recency, and semantic impact metrics. Integrated with EdgeLoRA, SemCache reduces memory cost up to 10.8× and achieves 2.11× latency speedup per query via efficient token reuse, offering a scalable, privacy-preserving solution for efficient LoRA-adapted LLM inference at the edge.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache OptimizationBorui Li, Yitao Wang, Haoran Ma, Ligeng Chen et al.ACL 2025 · 6 citations
- ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM ServingJiuchen Shi, Hang Zhang, Yixiao Wang, Quan Chen et al.HPCA 2026
- Federated Adaptive Fine-Tuning of Large Language Models with Heterogeneous Quantization and LoRAZhidong Gao, Zhenxiao Zhang, Yuanxiong Guo, Yanmin GongINFOCOM 2025 · 12 citations
- BSLoRA: Enhancing the Parameter Efficiency of LoRA with Intra-Layer and Inter-Layer SharingYuhua Zhou, Ruifeng Li, Changhai Zhou, Fei Yang et al.ICML 2025
- Don't Reinvent the Wheel, Just Realign the Spokes: Resource-Efficient Federated Fine-Tuning via Rank-Wise Expert AssemblyYebo Wu, Jingguang Li, Zhijiang Guo, Li LiICML 2026
