VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG
Junkyum Kim, Divya Mahajan
Abstract
Retrieval-Augmented Generation leverages vector similarity search to enhance large language models with up-to-date, external knowledge, enabling accurate and reliable responses. While CPU-only vector search incurs high latency on large, high-dimensional indices, co-locating the retriever and the LLM on the GPU leads to resource sharing that can create resource contention. Specifically, vector search is memory and I/O intensive, placing it in direct conflict with LLM inference, which demands memory for KV cache and compute for higher throughput. We present VectorLiterAG, a latency-aware RAG serving system that explicitly orchestrates data placement and execution across retrieval and inference to meet strict end-to-end SLOs. VectorLiterAG is driven by access-pattern analysis and performance estimation to regulate how retrieval variability can be mitigated and managed in the system with LLM inference, enabling SLO-compliant execution under skewed and dynamic workloads. By jointly modeling search latency and query hit-rate distributions, VectorLiterAG identifies an optimal index partitioning point across CPU and GPU that minimizes contention and stabilizes batching behavior, thereby maximizing sustained throughput under skewed access patterns. A low-overhead online index update mechanism allows VectorLiterag to continuously adapt to evolving request distributions, preserving batching efficiency and throughput as access patterns evolve. Our evaluations demonstrate that VecTORLITERAG consistently expands the range of SLO-compliant request rate across all tested configurations. Without increasing the generation latency or requiring additional hardware, VecTORLITERAG outperforms both naive and existing alternative frameworks, improving attainable SLO-bound throughput by up to.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0292e17-bb56-4748-b974-019356da157dBuilds on16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
Related papers
- Omnia: Efficient RAG Serving through Speculative SchedulingRongtian Fu, Shigang Li, Youxuan Xu, Tong Wu et al.HPDC 2026
- Accelerating Graph-Based RAG Retrieval via Locality-Aware Device-Cloud CollaborationYongheng Deng, Tianyuan Jiang, Zhenya Ma, Hao Wu et al.KDD 2026
- Hermes: Algorithm-System Co-design for Efficient Retrieval-Augmented Generation At-ScaleMichael Shen, Muhammad Umar, Kiwan Maeng, G. Edward Suh et al.ISCA 2025 · 5 citations
- DReX: Accurate and Scalable Dense Retrieval Acceleration via Algorithmic-Hardware CodesignDerrick Quinn, E. Ezgi Yücel, Martin Prammer, Zhenxing Fan et al.ISCA 2025 · 10 citations
- RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation ServingWenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso et al.ISCA 2025 · 16 citations
