Lune

HPDC2026Top-tier venue

Omnia: Efficient RAG Serving through Speculative Scheduling

Rongtian Fu, Shigang Li, Youxuan Xu, Tong Wu, Zhi Ma, Jinliang Shi

2026Year

Abstract

Retrieval-Augmented Generation (RAG) has emerged for enhancing Large Language Models (LLMs) by improving factual accuracy and mitigating hallucinations. A typical RAG pipeline executes in three cascaded stages: retrieval, reranking, and generation. The existing serving systems suffer from two critical system-level bottlenecks when applying to RAG serving: the first is the cumulative latency caused by rigid sequential dependencies between reranking and generation, and the second is the system saturation triggered by bursty, high fan-in reranking workloads.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 7ecc0f7a-2c94-45fb-8f01-4096573c792c

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines