Lune

ICDE2026Top-tier venue

EC-RAG: Towards Efficient Edge-Cloud Retrieval-Augmented Generation Systems

Liang Wang, Kai Wang, Ranjun Jia, Kai Lu, Jiguang Wan, Hao Huo, Yulong Zhai, Zhiyuan Liang, Di Wang

2026Year

Abstract

Retrieval-augmented generation (RAG) grounds the desired responses of large language models (LLMs) by connecting to external knowledge databases. However, deploying full-scale LLMs on resource-constrained edge servers is impractical. Instead, small language models (SLMs) enable efficient deployment at the edge, but their limited model capacity may compromise generation quality. In this paper, we propose EC-RAG, a novel edge-cloud RAG system that selectively decides whether the final response should be generated by the edge-based SLM or the API-accessed cloud-based LLM. Specifically, on top of standard RAG, EC-RAG integrates dynamic chunk pruning to retain only useful information for each query and achieves adaptive query routing based on the query complexity and the amount of retained information, balancing accuracy, latency, and API cost. Furthermore, we explore several system optimizations to improve SLM inference performance at the edge. Our evaluation across four datasets shows that EC-RAG incurs only a slight additional cost relative to the SLM-only RAG and closes much of the quality gap with the LLM-only RAG. EC-RAG delivers up to a 30% higher F1-score than the SLM-only RAG and reduces costs by 90% compared to the LLM-only RAG, while ensuring low average latency.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines