PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
Kuntai Du, Bowen Wang, Chen Zhang, Yiming Cheng, Qing Lan, Hejian Sang, Yihua Cheng, Jiayi Yao, Xiaoxuan Liu, Yifan Qiao, Ion Stoica, Junchen Jiang
Abstract
Besides typical generative applications, like ChatGPT, GitHub Copilot, and Cursor, we observe an emerging trend that LLMs are increasingly used in traditional discriminative tasks, such as recommendation, credit verification, and data labeling. The key characteristic of these emerging use cases is that the LLM generates only a single output token, rather than an arbitrarily long sequence of tokens. We refer to this as a prefill-only workload. However, since existing LLM engines assume arbitrary output lengths, they fail to leverage the unique properties of prefill-only workloads. In this paper, we present PrefillOnly, the first LLM inference engine that improves the inference throughput and latency by fully embracing the properties of prefill-only workloads. First, since it generates only one token, PrefillOnly only needs to store the KV cache of only the last computed layer, rather than of all layers. This drastically reduces the GPU memory footprint of LLM inference and allows handling long inputs without using solutions that reduce throughput, such as cross-GPU KV cache parallelization. Second, because the output length is fixed, rather than arbitrary, PrefillOnly can precisely determine the job completion time (JCT) of each prefill-only request before it starts. This enables efficient JCT-aware scheduling policies such as shortest prefill first. PrefillOnly can process up to 4× larger queries per second without inflating the average and P99 latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Bat: Efficient Generative Recommender Serving with Bipartite AttentionJie Sun, Shaohang Wang, Zimo Zhang, Zhengyu Liu et al.ASPLOS 2026 · 1 citation
- On-device Semantic Selection Made Low Latency and Memory Efficient with Monolithic ForwardingJiahao Zhou, Chengliang Lin, Dingji Li, Mingkai Dong et al.EuroSys 2026
- Libra: Effective yet Efficient Load Balancing for Large-scale MoE InferenceJaehoon Yang, Yushin Kim, Seokwon Moon, Yeonhong Park et al.ICLR 2026
- MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block RepresentationsJiaxiang Zou, Yonghao Chen, Ruilong WU, Xinyu ChenICML 2026
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
Related papers
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae et al.HPDC 2026
- Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance EstimationJingyu Liu, Beidi Chen, Ce ZhangICML 2025
- FastServe: Iteration-Level Preemptive Scheduling for Large Language Model InferenceBingyang Wu, Yinmin Zhong, Zili Zhang, Shengyu Liu et al.NSDI 2026 · 12 citations
- SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model TransformationAurick Qiao, Zhewei Yao, Samyam Rajbhandari, Yuxiong HeEMNLP 2025 · 1 citation
- Efficient LLM Scheduling by Learning to RankYichao Fu, Siqi Zhu, Runlong Su, Aurick Qiao et al.NeurIPS 2024 · 129 citations
