AdaRAG: Adaptive Optimization for Retrieval Augmented Generation with Multilevel Retrievers at the Edge
Tao Ouyang, Guihang Hong, Kongyange Zhao, Zhi Zhou, Weigang Wu, Zhaobiao Lv, Xu Chen
Abstract
Considering privacy concerns and real-time demands of popular large language models (LLMs), a shift towards edge-based LLM inference leverages edge clusters in proximity to provide low latency and secure responsiveness. To enhance the generation quality of LLMs, retrieval-augmented generation (RAG) can seamlessly integrate relevant external knowledge from local databases into LLMs without dedicated fine-tuning. However, this retrieval process can significantly contribute to overall latency, particularly in resource-constrained edge environments. To address this challenge, we introduce AdaRAG, tailored for edge-based RAG, leveraging multilevel (i.e., light and heavy) retrievers to facilitate adaptive retrieval granularity and efficient pipeline parallelism for retrieval and inference processes by fully exploiting heterogeneous edge resources (i.e., CPU and GPU). AdaRAG adaptively manages the heavy retrieval proportion and selected documents in augmented prompts, aiming to balance the long-term trade-off between overall generation quality and latency for dynamic user queries. Due to the inherent randomness of probabilistic LLM inference and highly dynamic queries at the edge, the underlying relations between the above decisions and performance feedback (i.e., end-to-end latency and accuracy) are difficult to obtain accurately a priori. Thus, we adopt bandit convex optimization to design a lightweight online algorithm, which utilizes real-time performance feedback to estimate the gradient information and optimize the retrieval and prompt decisions on the fly. Our rigorous theoretical analysis and extensive evaluations show our AdaRAG's superior performance. These promising results can boost the adoption of AdaRAG in future edge-based LLM applications.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get edcc2259-f4b2-4ac7-928b-e4b67c0d1f7bCited by top-tier papers2
- CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge ComputingGuihang Hong, Tao Ouyang, Kongyange Zhao, Zhi Zhou et al.RTSS 2025 · 4 citations
- FedMosaic: Federated Retrieval-Augmented Generation via Parametric AdaptersZhilin Liang, Yuxiang Wang, Zimu Zhou, Hainan Zhang et al.SIGIR 2026
Related papers
- EC-RAG: Towards Efficient Edge-Cloud Retrieval-Augmented Generation SystemsLiang Wang, Kai Wang, Ranjun Jia, Kai Lu et al.ICDE 2026
- SRAG: A Lightweight and Specialized Retrieval-augmented Generation System at the EdgeRuikun Luo, Zihan Xing, Lin Gu, Song Wu et al.SIGIR 2026
- METIS: Fast Quality-Aware RAG Systems with Configuration AdaptationSiddhant Ray, Rui Pan, Zhuohan Gu, Kuntai Du et al.SOSP 2025 · 3 citations
- SubGCache: Accelerating Graph-based RAG with Subgraph-level KV CacheQiuyu Zhu, Liang Zhang, Qianxiong Xu, Cheng Long et al.AAAI 2026 · 1 citation
- RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation ServingWenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso et al.ISCA 2025 · 16 citations
