Lune

INFOCOM2025顶会

AdaRAG: Adaptive Optimization for Retrieval Augmented Generation with Multilevel Retrievers at the Edge

Tao Ouyang, Guihang Hong, Kongyange Zhao, Zhi Zhou, Weigang Wu, Zhaobiao Lv, Xu Chen

2025年份
5被引次数
2顶会引用

摘要

Considering privacy concerns and real-time demands of popular large language models (LLMs), a shift towards edge-based LLM inference leverages edge clusters in proximity to provide low latency and secure responsiveness. To enhance the generation quality of LLMs, retrieval-augmented generation (RAG) can seamlessly integrate relevant external knowledge from local databases into LLMs without dedicated fine-tuning. However, this retrieval process can significantly contribute to overall latency, particularly in resource-constrained edge environments. To address this challenge, we introduce AdaRAG, tailored for edge-based RAG, leveraging multilevel (i.e., light and heavy) retrievers to facilitate adaptive retrieval granularity and efficient pipeline parallelism for retrieval and inference processes by fully exploiting heterogeneous edge resources (i.e., CPU and GPU). AdaRAG adaptively manages the heavy retrieval proportion and selected documents in augmented prompts, aiming to balance the long-term trade-off between overall generation quality and latency for dynamic user queries. Due to the inherent randomness of probabilistic LLM inference and highly dynamic queries at the edge, the underlying relations between the above decisions and performance feedback (i.e., end-to-end latency and accuracy) are difficult to obtain accurately a priori. Thus, we adopt bandit convex optimization to design a lightweight online algorithm, which utilizes real-time performance feedback to estimate the gradient information and optimize the retrieval and prompt decisions on the fly. Our rigorous theoretical analysis and extensive evaluations show our AdaRAG's superior performance. These promising results can boost the adoption of AdaRAG in future edge-based LLM applications.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖