MVR-cache: Optimizing Semantic Caching via Multi-Vector Retrieval and Learned Prompt Segmentation
Ali Noshad, Zishan Zheng, Yinjun Wu
Abstract
To reduce LLM costs and latency, semantic caching systems must accurately identify when a new prompt matches a cached one. Current methods often rely on simplistic similarity measures, which limit their effectiveness. We introduce MVR-cache, a novel semantic caching approach that significantly improves retrieval accuracy by integrating Multi-Vector Retrieval (MVR). MVR-cache is built upon a learnable segmentation model that intelligently splits prompts, enabling fine-grained similarity comparisons via MaxSim. We derive the model's training objective from a rigorous theoretical analysis. This can ensure that optimizing this objective directly maximizes cache hits under strict correctness constraints. To solve the resulting non-differentiable combinatorial optimization problem, we leverage a reinforcement learning-based training strategy with the theoretically grounded objectives as the reward. Experimental results on established benchmarks across diverse tasks confirm that in comparison to the state-of-the-art, MVR-cache consistently increases the cache hit rates by up to 37% while maintaining the same correctness guarantees. MVR-cache is available at https://github.com/PKU-SDS-lab/MVR-Cache
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20a025be-8ff0-48d0-be6b-163d48ca9300Builds on6
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Explicit Query Rewriting for Conversational Dense RetrievalHongjin Qian, Zhicheng DouEMNLP 2022 · 14 citations
- Threshold-Consistent Margin Loss for Open-World Deep Metric LearningQin Zhang, Linghan Xu, Jun Fang, Qingming Tang et al.ICLR 2024 · 10 citations
- CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector RetrievalMinghan Li, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal et al.ACL 2023 · 10 citations
- Generalized Pseudo-Relevance FeedbackYiteng Tu, Weihang Su, Yujia Zhou, Yiqun Liu et al.WWW 2026 · 2 citations
Related papers
- vCache: Verified Semantic Prompt CachingLuis Gaspar Schroeder, Aditya Desai, Alejandro Cuadron, Kyle Chu et al.ICLR 2026 · 15 citations
- Generative Caching for Structurally Similar Prompts and ResponsesSarthak Chakraborty, Suman Nath, Xuchao Zhang, Chetan Bansal et al.NeurIPS 2025 · 5 citations
- CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for Accelerating LLM ServingYang Liu, Yunfei Gu, Liqiang Zhang, Chentao Wu et al.FAST 2026 · 14 citations
- POQD: Performance-Oriented Query Decomposer for Multi-vector retrievalYaoyang Liu, Junlin Li, Yinjun Wu, Zhen ChenICML 2025
- Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge CachingChaoyi Ruan, Chao Bi, Kaiwen Zheng, Ziji Shi et al.NSDI 2026 · 6 citations
