Omnia: Efficient RAG Serving through Speculative Scheduling
Rongtian Fu, Shigang Li, Youxuan Xu, Tong Wu, Zhi Ma, Jinliang Shi
Abstract
Retrieval-Augmented Generation (RAG) has emerged for enhancing Large Language Models (LLMs) by improving factual accuracy and mitigating hallucinations. A typical RAG pipeline executes in three cascaded stages: retrieval, reranking, and generation. The existing serving systems suffer from two critical system-level bottlenecks when applying to RAG serving: the first is the cumulative latency caused by rigid sequential dependencies between reranking and generation, and the second is the system saturation triggered by bursty, high fan-in reranking workloads.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7ecc0f7a-2c94-45fb-8f01-4096573c792cRelated papers
- METIS: Fast Quality-Aware RAG Systems with Configuration AdaptationSiddhant Ray, Rui Pan, Zhuohan Gu, Kuntai Du et al.SOSP 2025 · 3 citations
- RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation ServingWenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso et al.ISCA 2025 · 16 citations
- VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAGJunkyum Kim, Divya MahajanHPCA 2026 · 2 citations
- REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage ProcessingKangqi Chen, Rakesh Nadig, Manos Frouzakis, Nika Mansouri-Ghiasi et al.ISCA 2025 · 14 citations
- Improving Retrieval-Augmented Generation through Multi-Agent Reinforcement LearningYiqun Chen, Lingyong Yan, Weiwei Sun, Xinyu Ma et al.NeurIPS 2025 · 47 citations
