Lune

INFOCOM2026Top-tier venue

SemCache: Semantic-Aware Cache Sharing for Efficient Multi-User LoRA-Adapted LLM Inference at the Edge

Tao Ren, Yiming Yao, Zheyuan Hu, Jianwei Niu

2026Year
1Citations

Abstract

Large language models (LLMs) enable personalized applications through parameter-efficient techniques like Low-Rank Adaptation (LoRA). Deploying LoRA-adapted LLMs faces a critical dilemma: resource-constrained user devices (UDs) are hard to host full-scale LLMs, while centralized edge deployments of all LoRA adapters incur huge storage and privacy risks. We propose EdgeLoRA, a collaborative framework where UDs retain private LoRA adapters locally while offloading base model computations to edge servers (ESs), enhancing privacy and reducing storage overhead. However, EdgeLoRA introduces latency inefficiency from multi-round ES-UD interactions and memory overhead from redundant per-user key-value (KV) caching. To address this, we design SemCache, a semantic-aware caching mechanism that exploits hierarchical redundancies. SemCache first clusters user queries into coarse-grained intents (e.g., health, economic) via lightweight textual encoders, then identifies reusable fine-grained subsequences (e.g., "feel pain in") within clusters. A hybrid-indexed global cache (cluster index + token position) enables cross-user KV reuse, while an adaptive policy optimizes cache utility via frequency, recency, and semantic impact metrics. Integrated with EdgeLoRA, SemCache reduces memory cost up to 10.8× and achieves 2.11× latency speedup per query via efficient token reuse, offering a scalable, privacy-preserving solution for efficient LoRA-adapted LLM inference at the edge.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines