Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
Chaoyi Ruan, Chao Bi, Kaiwen Zheng, Ziji Shi, Xinyi Wan, Jialin Li
Abstract
Large Language Model (LLM) agents tackle data-intensive tasks such as deep research and code generation. However, their effectiveness depends on frequent interactions with knowledge sources across remote clouds or regions. Such interactions can create non-trivial latency and cost bottlenecks. Existing caching solutions focus on exact-match queries, limiting their effectiveness for semantic knowledge reuse.
To address this challenge, we introduce Cortex, a novel cross-region knowledge caching architecture for LLM agents. At its core are two abstractions: Semantic Element (SE) and Semantic Retrieval Index (Seri). A semantic element captures the semantic embedding representation of an LLM query together with performance-aware metadata such as latency, cost, and staticity. Seri then provides a two-stage retrieval pipeline: a vector similar index with semantic embedding for fast candidate selection and a lightweight LLM-powered semantic judge for precise validation. Atop these primitives, Cortex builds a new cache interface that includes a new semanticaware cache hit definition, a cost-efficient eviction policy, and proactive prefetching. To reduce overhead, Cortex co-locates the smaller LLM judge with the main LLM using adaptive scheduling and resource sharing. Our evaluation demonstrates that Cortex delivers substantial performance improvements without compromising correctness. On representative search workloads, Cortex achieves up to a 3.6× increase in throughput by maintaining cache hit rates of over 85% while preserving accuracy virtually identical to non-cached baselines. Cortex also improves throughput for coding tasks by 20%, showcasing its versatility across diverse agentic workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65804b4f-544d-4ceb-bf0e-c2a2a8740d26Builds on10
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- A-Mem: Agentic Memory for LLM AgentsWujiang Xu, Zujie Liang, Kai Mei, Hang Gao et al.NeurIPS 2025 · 1,138 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye et al.AAAI 2024 · 394 citations
- Parrot: Efficient Serving of LLM-based Applications with Semantic VariableChaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang et al.OSDI 2024 · 112 citations
Related papers
- Demystifying and Enhancing the Efficiency of Large Language Model Based Search AgentsTiannuo Yang, Zebin Yao, Bowen Jin, Lixiao Cui et al.ICLR 2026 · 9 citations
- Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM AgentsQizheng Zhang, Michael Wornow, Kunle OlukotunNeurIPS 2025 · 27 citations
- Generative Caching for Structurally Similar Prompts and ResponsesSarthak Chakraborty, Suman Nath, Xuchao Zhang, Chetan Bansal et al.NeurIPS 2025 · 5 citations
- CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for Accelerating LLM ServingYang Liu, Yunfei Gu, Liqiang Zhang, Chentao Wu et al.FAST 2026 · 14 citations
- KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent WorkflowsZaifeng Pan, Ajjkumar Patel, Yipeng Shen, Zhengding Hu et al.NeurIPS 2025 · 77 citations
