Omnia: Efficient RAG Serving through Speculative Scheduling
Rongtian Fu, Shigang Li, Youxuan Xu, Tong Wu, Zhi Ma, Jinliang Shi
2026年份
摘要
Retrieval-Augmented Generation (RAG) has emerged for enhancing Large Language Models (LLMs) by improving factual accuracy and mitigating hallucinations. A typical RAG pipeline executes in three cascaded stages: retrieval, reranking, and generation. The existing serving systems suffer from two critical system-level bottlenecks when applying to RAG serving: the first is the cumulative latency caused by rigid sequential dependencies between reranking and generation, and the second is the system saturation triggered by bursty, high fan-in reranking workloads.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- METIS: Fast Quality-Aware RAG Systems with Configuration AdaptationSiddhant Ray, Rui Pan, Zhuohan Gu, Kuntai Du 等SOSP 2025 · 被引用 3 次
- RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation ServingWenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso 等ISCA 2025 · 被引用 16 次
- VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAGJunkyum Kim, Divya MahajanHPCA 2026 · 被引用 2 次
- REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage ProcessingKangqi Chen, Rakesh Nadig, Manos Frouzakis, Nika Mansouri-Ghiasi 等ISCA 2025 · 被引用 14 次
- Improving Retrieval-Augmented Generation through Multi-Agent Reinforcement LearningYiqun Chen, Lingyong Yan, Weiwei Sun, Xinyu Ma 等NeurIPS 2025 · 被引用 47 次
