SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning
Lingkun Long, Rubing Yang, Yushi Huang, Desheng Hui, Ao Zhou, Jianlei Yang
摘要
Long-context inference for Large Language Models (LLMs) is heavily limited by high computational demands. While several existing methods optimize attention computation, they still process the full set of hidden states at each layer, limiting overall efficiency. In this work, we propose SlimInfer, an innovative framework that aims to accelerate inference by directly pruning less critical prompt tokens during the forward pass. Our key insight is an information diffusion phenomenon: As information from critical tokens propagates through layers, it becomes distributed across the entire sequence. This diffusion process suggests that LLMs can maintain their semantic integrity when excessive tokens, even including these critical ones, are pruned in hidden states. Motivated by this, SlimInfer introduces a dynamic fine-grained pruning mechanism that accurately removes redundant tokens of hidden state at intermediate layers. This layer-wise pruning naturally enables an asynchronous KV cache manager that prefetches required token blocks without complex predictors, reducing both memory usage and I/O costs. Extensive experiments show that SlimInfer can achieve up to 2.53× timeto-first-token (TTFT) speedup and 1.88× end-to-end latency reduction for LLaMA3.1-8B-Instruct on a single RTX 4090, without sacrificing performance on LongBench. Our code will be available at https://github.com/Longxmas/SlimInfer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in InferenceHao Zhang, Mengsi Lyu, Zhuo Chen, Yulong Ao 等ACL 2026 · 被引用 9 次
- Stability Implies Redundancy: Delta Attention Selective Halting for Efficient Long-Context PrefillingYujie Chen, Tailai Chen, Yifeng Gao, Zoe Wanying He 等ACL 2026 · 被引用 1 次
- On-device Semantic Selection Made Low Latency and Memory Efficient with Monolithic ForwardingJiahao Zhou, Chengliang Lin, Dingji Li, Mingkai Dong 等EuroSys 2026
它引用的顶会 Paper17
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
相关 Paper
- HShare: Fast LLM Decoding by Hierarchical Key-Value SharingHuaijin Wu, Lianqiang Li, Hantao Huang, Tu Yi 等ICLR 2025
- Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance EstimationJingyu Liu, Beidi Chen, Ce ZhangICML 2025
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 被引用 1 次
- Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache EvictionYuerong Song, Xiaoran Liu, Ruixiao Li, Zhigeng Liu 等AAAI 2026 · 被引用 43 次
- TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache SelectionWei Wu, Zhuoshi Pan, Kun Fu, Chao Wang 等EMNLP 2025 · 被引用 2 次
