Accelerating Iterative Retrieval-augmented Language Model Serving with Speculation
Zhihao Zhang, Alan Zhu, Lijie Yang, Yihua Xu, Lanting Li, Phitchaya Mangpo Phothilimthana, Zhihao Jia
摘要
This paper introduces RaLMSpec, a framework that accelerates iterative retrieval-augmented language model (RaLM) with speculative retrieval and batched verification. RaLMSpec further introduces several important systems optimizations, including prefetching, optimal speculation stride scheduler, and asynchronous verification. The combination of these techniques allows RaLM-SPec to significantly outperform existing systems. For document-level iterative RaLM serving, evaluation over three LLMs on four QA datasets shows that RaLMSpec improves over existing approaches by 1.75-2.39×, 1.04-1.39×, and 1.31-1.77× when the retriever is an exact dense retriever, approximate dense retriever, and sparse retriever respectively. For token-level iterative RaLM (KNN-LM) serving, RaLMSpec is up to 7.59× and 2.45× faster than existing methods for exact dense and approximate dense retrievers, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- METIS: Fast Quality-Aware RAG Systems with Configuration AdaptationSiddhant Ray, Rui Pan, Zhuohan Gu, Kuntai Du 等SOSP 2025 · 被引用 3 次
- Mnemosyne: Accelerating Multi-Hop Question Answering via Cache Hit Order FittingHaizhou Du, Jiujiu Li, Dongyang Li, Luobin Huang 等AAAI 2026
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
相关 Paper
- Speculative RAG: Enhancing Retrieval Augmented Generation through DraftingZilong Wang, Zifeng Wang, Long T. Le, Huaixiu Steven Zheng 等ICLR 2025 · 被引用 7 次
- RAPID: Long-Context Inference with Retrieval-Augmented Speculative DecodingGuanzheng Chen, Qilong Feng, Jinjie Ni, Xin Li 等ICML 2025
- FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative SamplingWeilin Zhao, Tengyu Pan, Xu Han, Yudi Zhang 等ACL 2025 · 被引用 14 次
- Efficiency Unleashed: Inference Acceleration for LLM-based Recommender Systems with Speculative DecodingYunjia Xi, Hangyu Wang, Bo Chen, Jianghao Lin 等SIGIR 2025 · 被引用 5 次
- Accelerating Retrieval Augmented Language Model via PIM and PNM IntegrationJe-Woo Jang, Junyong Oh, Youngbae Kong, Jae-Youn Hong 等MICRO 2025 · 被引用 5 次
