VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG
Junkyum Kim, Divya Mahajan
摘要
Retrieval-Augmented Generation leverages vector similarity search to enhance large language models with up-to-date, external knowledge, enabling accurate and reliable responses. While CPU-only vector search incurs high latency on large, high-dimensional indices, co-locating the retriever and the LLM on the GPU leads to resource sharing that can create resource contention. Specifically, vector search is memory and I/O intensive, placing it in direct conflict with LLM inference, which demands memory for KV cache and compute for higher throughput. We present VectorLiterAG, a latency-aware RAG serving system that explicitly orchestrates data placement and execution across retrieval and inference to meet strict end-to-end SLOs. VectorLiterAG is driven by access-pattern analysis and performance estimation to regulate how retrieval variability can be mitigated and managed in the system with LLM inference, enabling SLO-compliant execution under skewed and dynamic workloads. By jointly modeling search latency and query hit-rate distributions, VectorLiterAG identifies an optimal index partitioning point across CPU and GPU that minimizes contention and stabilizes batching behavior, thereby maximizing sustained throughput under skewed access patterns. A low-overhead online index update mechanism allows VectorLiterag to continuously adapt to evolving request distributions, preserving batching efficiency and throughput as access patterns evolve. Our evaluations demonstrate that VecTORLITERAG consistently expands the range of SLO-compliant request rate across all tested configurations. Without increasing the generation latency or requiring additional hardware, VecTORLITERAG outperforms both naive and existing alternative frameworks, improving attainable SLO-bound throughput by up to.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu 等OSDI 2024 · 被引用 646 次
相关 Paper
- Omnia: Efficient RAG Serving through Speculative SchedulingRongtian Fu, Shigang Li, Youxuan Xu, Tong Wu 等HPDC 2026
- Accelerating Graph-Based RAG Retrieval via Locality-Aware Device-Cloud CollaborationYongheng Deng, Tianyuan Jiang, Zhenya Ma, Hao Wu 等KDD 2026
- Hermes: Algorithm-System Co-design for Efficient Retrieval-Augmented Generation At-ScaleMichael Shen, Muhammad Umar, Kiwan Maeng, G. Edward Suh 等ISCA 2025 · 被引用 5 次
- DReX: Accurate and Scalable Dense Retrieval Acceleration via Algorithmic-Hardware CodesignDerrick Quinn, E. Ezgi Yücel, Martin Prammer, Zhenxing Fan 等ISCA 2025 · 被引用 10 次
- RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation ServingWenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso 等ISCA 2025 · 被引用 16 次
