Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding
Jinglin Chen, Qiwei Li, Zuchao Li, Baoyuan Qi, Guoming Liu, Haojun Ai, Hai Zhao, Ping Wang
Abstract
As a crucial method in prompt engineering, In-Context Learning (ICL) enhances the generalization and knowledge utilization capabilities of Large Language Models (LLMs) (Dong et al., 2024) . However, the lengthy retrieved contexts and limited token throughput in autoregressive models significantly constrain reasoning speed. To address this challenge, we propose N-Gram Trie Speculative Decoding, a novel approach that leverages the overlap between context and model output. This method constructs an n-gram trie from the context to generate drafts, accelerating token generation for LLMs. We evaluate our approach on summarization, Retrieval-Augmented Generation (RAG), and contextbased Question Answering (QA) tasks. Experimental results on Vicuna-7B, Llama2-7B-Chat, and Llama3-8B-Instruct demonstrate substantial speed improvements without compromising accuracy. Compared with various strong baselines, our method achieves the highest mean speedup, showcasing its effectiveness and efficiency. Our implement code is available here: https://github.com/mrlife219/Ngram-Trie .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba8eb760-6f5c-4da7-84ef-6bfa6b16b886Cited by top-tier papers2
- Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch ScenariosLuohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi et al.AAAI 2026 · 1 citation
- From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic HorizonsXiangyu Ma, Teng Xiao, Zuchao Li, Lefei ZhangACL 2026
Builds on11
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng et al.ICML 2024 · 669 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 424 citations
Related papers
- RAPID: Long-Context Inference with Retrieval-Augmented Speculative DecodingGuanzheng Chen, Qilong Feng, Jinjie Ni, Xin Li et al.ICML 2025
- SAM Decoding: Speculative Decoding via Suffix AutomatonYuxuan Hu, Ke Wang, Xiaokang Zhang, Fanjin Zhang et al.ACL 2025
- An Empirical Study of Speculative Decoding on Software Engineering TasksYijia Li, Junkai Chen, Xing Hu, Xin XiaISSTA 2026
- Speculative Decoding with CTC-based Draft Model for LLM Inference AccelerationZhuofan Wen, Shangtong Gui, Yang FengNeurIPS 2024 · 19 citations
- GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative DecodingCunxiao Du, Jing Jiang, Yuanchen Xu, Jiawei Wu et al.ICML 2024 · 72 citations
