SemCache: Semantic-Aware Cache Sharing for Efficient Multi-User LoRA-Adapted LLM Inference at the Edge
Tao Ren, Yiming Yao, Zheyuan Hu, Jianwei Niu
摘要
Large language models (LLMs) enable personalized applications through parameter-efficient techniques like Low-Rank Adaptation (LoRA). Deploying LoRA-adapted LLMs faces a critical dilemma: resource-constrained user devices (UDs) are hard to host full-scale LLMs, while centralized edge deployments of all LoRA adapters incur huge storage and privacy risks. We propose EdgeLoRA, a collaborative framework where UDs retain private LoRA adapters locally while offloading base model computations to edge servers (ESs), enhancing privacy and reducing storage overhead. However, EdgeLoRA introduces latency inefficiency from multi-round ES-UD interactions and memory overhead from redundant per-user key-value (KV) caching. To address this, we design SemCache, a semantic-aware caching mechanism that exploits hierarchical redundancies. SemCache first clusters user queries into coarse-grained intents (e.g., health, economic) via lightweight textual encoders, then identifies reusable fine-grained subsequences (e.g., "feel pain in") within clusters. A hybrid-indexed global cache (cluster index + token position) enables cross-user KV reuse, while an adaptive policy optimizes cache utility via frequency, recency, and semantic impact metrics. Integrated with EdgeLoRA, SemCache reduces memory cost up to 10.8× and achieves 2.11× latency speedup per query via efficient token reuse, offering a scalable, privacy-preserving solution for efficient LoRA-adapted LLM inference at the edge.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache OptimizationBorui Li, Yitao Wang, Haoran Ma, Ligeng Chen 等ACL 2025 · 被引用 6 次
- ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM ServingJiuchen Shi, Hang Zhang, Yixiao Wang, Quan Chen 等HPCA 2026
- Federated Adaptive Fine-Tuning of Large Language Models with Heterogeneous Quantization and LoRAZhidong Gao, Zhenxiao Zhang, Yuanxiong Guo, Yanmin GongINFOCOM 2025 · 被引用 12 次
- BSLoRA: Enhancing the Parameter Efficiency of LoRA with Intra-Layer and Inter-Layer SharingYuhua Zhou, Ruifeng Li, Changhai Zhou, Fei Yang 等ICML 2025
- Don't Reinvent the Wheel, Just Realign the Spokes: Resource-Efficient Federated Fine-Tuning via Rank-Wise Expert AssemblyYebo Wu, Jingguang Li, Zhijiang Guo, Li LiICML 2026
