MVR-cache: Optimizing Semantic Caching via Multi-Vector Retrieval and Learned Prompt Segmentation
Ali Noshad, Zishan Zheng, Yinjun Wu
摘要
To reduce LLM costs and latency, semantic caching systems must accurately identify when a new prompt matches a cached one. Current methods often rely on simplistic similarity measures, which limit their effectiveness. We introduce MVR-cache, a novel semantic caching approach that significantly improves retrieval accuracy by integrating Multi-Vector Retrieval (MVR). MVR-cache is built upon a learnable segmentation model that intelligently splits prompts, enabling fine-grained similarity comparisons via MaxSim. We derive the model's training objective from a rigorous theoretical analysis. This can ensure that optimizing this objective directly maximizes cache hits under strict correctness constraints. To solve the resulting non-differentiable combinatorial optimization problem, we leverage a reinforcement learning-based training strategy with the theoretically grounded objectives as the reward. Experimental results on established benchmarks across diverse tasks confirm that in comparison to the state-of-the-art, MVR-cache consistently increases the cache hit rates by up to 37% while maintaining the same correctness guarantees. MVR-cache is available at https://github.com/PKU-SDS-lab/MVR-Cache
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Explicit Query Rewriting for Conversational Dense RetrievalHongjin Qian, Zhicheng DouEMNLP 2022 · 被引用 14 次
- Threshold-Consistent Margin Loss for Open-World Deep Metric LearningQin Zhang, Linghan Xu, Jun Fang, Qingming Tang 等ICLR 2024 · 被引用 10 次
- CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector RetrievalMinghan Li, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal 等ACL 2023 · 被引用 10 次
- Generalized Pseudo-Relevance FeedbackYiteng Tu, Weihang Su, Yujia Zhou, Yiqun Liu 等WWW 2026 · 被引用 2 次
相关 Paper
- vCache: Verified Semantic Prompt CachingLuis Gaspar Schroeder, Aditya Desai, Alejandro Cuadron, Kyle Chu 等ICLR 2026 · 被引用 15 次
- Generative Caching for Structurally Similar Prompts and ResponsesSarthak Chakraborty, Suman Nath, Xuchao Zhang, Chetan Bansal 等NeurIPS 2025 · 被引用 5 次
- CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for Accelerating LLM ServingYang Liu, Yunfei Gu, Liqiang Zhang, Chentao Wu 等FAST 2026 · 被引用 14 次
- POQD: Performance-Oriented Query Decomposer for Multi-vector retrievalYaoyang Liu, Junlin Li, Yinjun Wu, Zhen ChenICML 2025
- Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge CachingChaoyi Ruan, Chao Bi, Kaiwen Zheng, Ziji Shi 等NSDI 2026 · 被引用 6 次
