PipeRAG: Fast Retrieval-Augmented Generation via Adaptive Pipeline Parallelism
Wenqi Jiang, Shuai Zhang, Boran Han, Jie Wang, Bernie Wang, Tim Kraska
摘要
Retrieval-augmented generation (RAG) can enhance the generation quality of large language models (LLMs) by incorporating external token databases. However, retrievals from large databases can constitute a substantial portion of the overall generation time, particularly when retrievals are periodically performed to align the retrieved content with the latest states of generation. In this paper, we introduce PipeRAG, a novel algorithm-system co-design approach to reduce generation latency and enhance generation quality. PipeRAG integrates (1) pipeline parallelism to enable concurrent retrieval and generation processes, (2) flexible retrieval intervals to maximize the efficiency of pipeline parallelism, and (3) a performance model to automatically balance retrieval quality and latency based on the generation states and underlying hardware. Our evaluation shows that, by combining the three aforementioned methods, PipeRAG achieves up to 2.6× speedup in end-to-end generation latency while improving generation quality. These promising results showcase the effectiveness of co-designing algorithms with underlying systems, paving the way for the adoption of PipeRAG in future RAG systems.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization TraversalWenqi Jiang, Hang Hu, Torsten Hoefler, Gustavo AlonsoVLDB 2025 · 被引用 10 次
- Demystifying and Enhancing the Efficiency of Large Language Model Based Search AgentsTiannuo Yang, Zebin Yao, Bowen Jin, Lixiao Cui 等ICLR 2026 · 被引用 9 次
- CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge ComputingGuihang Hong, Tao Ouyang, Kongyange Zhao, Zhi Zhou 等RTSS 2025 · 被引用 4 次
- TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized AssistantRongpei Hong, Jian Lang, Ting Zhong, Yong Wang 等KDD 2026
- Predictive Prefetching for Retrieval-Augmented GenerationWuyang Zhang, Shichao PeiICML 2026
相关 Paper
- Hermes: Algorithm-System Co-design for Efficient Retrieval-Augmented Generation At-ScaleMichael Shen, Muhammad Umar, Kiwan Maeng, G. Edward Suh 等ISCA 2025 · 被引用 5 次
- METIS: Fast Quality-Aware RAG Systems with Configuration AdaptationSiddhant Ray, Rui Pan, Zhuohan Gu, Kuntai Du 等SOSP 2025 · 被引用 3 次
- AquaPipe: A Quality-Aware Pipeline for Knowledge Retrieval and Large Language ModelsRunjie Yu, Weizhou Huang, Shuhan Bai, Jian Zhou 等SIGMOD 2025 · 被引用 6 次
- TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked TextSongshuo Lu, Hua Wang, Yutian Rong, Zhi Chen 等EMNLP 2025 · 被引用 2 次
- RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation ServingWenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso 等ISCA 2025 · 被引用 16 次
