BRIDGE: Block-Wise Speculative Coordination for Cloud-Edge Retrieval-Augmented Generation
Yuting Li, Shaoyuan Huang, Xiangqi Liu, Yunfeng Zhao, Xiaofei Wang
摘要
Retrieval-augmented generation (RAG) improves factuality by conditioning LLMs on retrieved evidence, yet real-world knowledge is often split across tiers: cloud-based RAG can exploit large public corpora, whereas edge-based RAG is the natural place to access private, user-specific stores. This raises a key question: how can one jointly leverage cloud and edge knowledge without transferring sensitive edge data or incurring prohibitive key-value (KV) recomputation. A direct output-level aggregation requires per-token synchronization, causing severe stalls under heterogeneous decoding speeds, and token-wise speculation still suffers from frequent communication and rollback overhead, limiting efficiency and usability. To address these challenges, we present BRIDGE, a Block-wise speculative RAG framework for Inter-database Distributed GEneration. BRIDGE enables parallel cloud--edge retrieval and drafting, without exposing private data or recomputing KV states. It introduces block-wise transmission and verification to amortize communication cost and improve acceptance efficiency, thereby reducing rollback waste. BRIDGE further employs an adaptive block-length strategy to balance the benefits of larger blocks against their verification overhead. Across public-private RAG benchmarks, model families, and network regimes, BRIDGE significantly improves question-answer relevance and personalization, while reducing end-to-end latency by up to 79.1%.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SRAG: A Lightweight and Specialized Retrieval-augmented Generation System at the EdgeRuikun Luo, Zihan Xing, Lin Gu, Song Wu 等SIGIR 2026
- AdaRAG: Adaptive Optimization for Retrieval Augmented Generation with Multilevel Retrievers at the EdgeTao Ouyang, Guihang Hong, Kongyange Zhao, Zhi Zhou 等INFOCOM 2025 · 被引用 5 次
- Bridging the Preference Gap between Retrievers and LLMsZixuan Ke, Weize Kong, Cheng Li, Mingyang Zhang 等ACL 2024 · 被引用 8 次
- Speculative RAG: Enhancing Retrieval Augmented Generation through DraftingZilong Wang, Zifeng Wang, Long T. Le, Huaixiu Steven Zheng 等ICLR 2025 · 被引用 7 次
- EC-RAG: Towards Efficient Edge-Cloud Retrieval-Augmented Generation SystemsLiang Wang, Kai Wang, Ranjun Jia, Kai Lu 等ICDE 2026
