Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
Jingyu Liu, Beidi Chen, Ce Zhang
Abstract
Improving time-to-first-token (TTFT) is an essentially important objective in modern large language model (LLM) inference engines. Optimizing TTFT directly results in higher maximal QPS and meets the requirements of many critical applications. However, boosting TTFT is notoriously challenging since it is computebounded and the performance bottleneck shifts from the self-attention to the MLP part. We present SPECPREFILL 1 , a training free framework that accelerates the inference TTFT for both long and medium context queries based on the following insight: LLMs are generalized enough to preserve the quality given only a carefully chosen subset of prompt tokens. At its core, SPECPREFILL leverages a lightweight model to speculate locally important tokens based on the context. These tokens, along with the necessary positional information, are then sent to the main model for processing. We evaluate SPECPRE-FILL with a diverse set of tasks, followed by a comprehensive benchmarking of performance improvement both in a real end-to-end setting and ablation studies. SPECPREFILL manages to serve Llama-3.1-405B-Instruct-FP8 with up to 7× maximal end-to-end QPS on real downstream tasks and 7.66× TTFT improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53043f05-8f22-4239-9126-ac3e10aaf2f0Cited by top-tier papers7
- Draft-based Approximate Inference for LLMsKevin Galim, Ethan Ewer, Wonjun Kang, Minjae Lee et al.ICLR 2026 · 5 citations
- Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM ServingZongze Li, Jingyu Liu, Zach Xu, Yineng Zhang et al.ICML 2026 · 4 citations
- SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM PrefillingXiaodong Ji, Hailin Zhang, Fangcheng Fu, Bin CuiICML 2026 · 3 citations
- When to Think, When to Speak: Learning Disclosure Policies for LLM ReasoningJiaqi Wei, Xuehang Guo, Pengfei Yu, Xiang Zhang et al.ICML 2026 · 2 citations
- SpecCache: Speculative KV Cache Reuse for Efficient RAG ServingZijian Wen, Tao Zhang, Shuangwu Chen, Shenghao Ye et al.ACL 2026
Builds on19
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
Related papers
- SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token PruningLingkun Long, Rubing Yang, Yushi Huang, Desheng Hui et al.AAAI 2026 · 8 citations
- SPECTRA: Faster Large Language Model Inference with Optimized Internal and External SpeculationNguyen-Khang Le, Truong Dinh Do, Le-Minh NguyenACL 2025
- SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model TransformationAurick Qiao, Zhewei Yao, Samyam Rajbhandari, Yuxiong HeEMNLP 2025 · 1 citation
- BOLT: Fewer Tokens but More Performance Retention for Efficient Vision-Language Models InferenceJiahua Bao, Siyao Cheng, Jiaxing Du, Changjiang He et al.ACM MM 2025
- Efficient Training-Free Multi-Token Prediction via Embedding-Space ProbingRaghavv Goel, Mukul Gagrani, Mingu Lee, Christopher LottICML 2026
