FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
Xunhao Lai, Jianqiao Lu, Yao Luo, Yiyuan Ma, Xun Zhou
摘要
Large language models (LLMs) encounter computational challenges during long-sequence inference, especially in the attention pre-filling phase, where the complexity grows quadratically with the prompt length. Previous efforts to mitigate these challenges have relied on fixed sparse attention patterns or identifying sparse attention patterns based on limited cases. However, these methods lacked the flexibility to efficiently adapt to varying input demands. In this paper, we introduce FlexPrefill, a Flexible sparse Pre-filling mechanism that dynamically adjusts sparse attention patterns and computational budget in real-time to meet the specific requirements of each input and attention head. The flexibility of our method is demonstrated through two key innovations: 1) Query-Aware Sparse Pattern Determination: By measuring Jensen-Shannon divergence, this component adaptively switches between query-specific diverse attention patterns and predefined attention patterns. 2) Cumulative-Attention Based Index Selection: This component dynamically selects query-key indexes to be computed based on different attention patterns, ensuring the sum of attention scores meets a predefined threshold.FlexPrefill adaptively optimizes the sparse pattern and sparse ratio of each attention head based on the prompt, enhancing efficiency in long-sequence inference tasks. Experimental results show significant improvements in both speed and accuracy over prior methods, providing a more flexible and efficient solution for LLM inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper57
- Neural Attention SearchDifan Deng, Marius LindauerNeurIPS 2025 · 被引用 431 次
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware PermutationShuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li 等NeurIPS 2025 · 被引用 114 次
- InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long AdaptationWeilin Zhao, Zihan Zhou, Zhou Su, Chaojun Xiao 等ICLR 2026 · 被引用 32 次
- Block-Sparse Global Attention for Efficient Multi-View Geometry TransformersChung-Shien Brian Wang, Christian Schmidt, Jens Piekenbrinck, Bastian LeibeCVPR 2026 · 被引用 22 次
- Flow Caching for Autoregressive Video GenerationYuexiao Ma, Xuzhe Zheng, Jing Xu, Xiwei Xu 等ICLR 2026 · 被引用 20 次
它引用的顶会 Paper31
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
相关 Paper
- Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient TransformersZecheng Tang, Quantong Qiu, Yi Yang, Zhiyi Hong 等ICML 2026 · 被引用 4 次
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu 等NeurIPS 2024 · 被引用 479 次
- ProxyAttn: Guided Sparse Attention via Representative HeadsYixuan Wang, Huang He, Siqi Bao, Hua Wu 等ICLR 2026 · 被引用 9 次
- AnchorAttention: Difference-Aware Sparse Attention with Stripe GranularityYu Zhang, Dong Guo, Fang Wu, Guoliang Zhu 等EMNLP 2025
- SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM PrefillingXiaodong Ji, Hailin Zhang, Fangcheng Fu, Bin CuiICML 2026 · 被引用 3 次
