IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs
Yuzhen Mao, Martin Ester, Ke Li
摘要
One limitation of existing Transformer-based models is that they cannot handle very long sequences as input since their self-attention operations exhibit quadratic time and space complexity. This problem becomes especially acute when Transformers are deployed on hardware platforms equipped only with CPUs. To address this issue, we propose a novel method for accelerating self-attention at inference time that works with pretrained Transformer models out-of-the-box without requiring retraining. We experiment using our method to accelerate various long-sequence Transformers, including a leading LLaMA 2-based LLM, on various benchmarks and demonstrate a speedup of 2.73 × -7.63× while retaining 98.6% -99.6% of the accuracy of the original pretrained models. The code is available on our project website at https://yuzhenmao.github.io/IceFormer/ . INTRODUCTION Transformers (Vaswani et al., 2017) have powered incredible advances in NLP, as exemplified by large language models (LLMs) such as GPT-4 and LLaMA 2. Increasingly LLMs are applied to exceptionally long input sequences, which enables many exciting applications such as long-form content creation, extended conversations, and large document search and analysis (OpenAI, 2023; Anthropic, 2023). While LLMs can be feasibly trained with expensive hardware accelerators (e.g. GPUs), they need to be deployed on commodity devices, which may only be equipped with CPUs. However, it is currently challenging to deploy LLMs on CPUs due to their high computation cost (Dice & Kogan, 2021) . A significant computational bottleneck arises from the self-attention mechanism that is integral to Transformers -both time and space complexity are quadratic in the sequence length. This problem is exacerbated in the context of LLMs, which are often used on very long sequences. To handle long input sequences, there has been substantial research into reducing the quadratic time complexity of self-attention -these methods are collectively known as efficient Transformers. However, many do not meet the needs of LLMs and are therefore difficult to apply to LLMs. An ideal acceleration method for LLMs should satisfy four criteria: (1) No retraining -the method should not require the model to be retrained, given the enormous computational expense of training LLMs; (2) Generality -the method should be applicable to a variety of LLMs, rather than just those trained with particular constraints built-in; (3) High accuracy -the method should not introduce large approximation errors, since LLMs have many attention layers and so errors from earlier layers can compound; (4) Fast inference -the method should achieve fast test-time performance. Satisfying all these criteria simultaneously is difficult, and to our knowledge no existing methods can do so. For example, Transformers with fixed attention patterns, e.g., Longformer (Beltagy et al., 2020) , require retraining the model before they can be used. Reformer (Nikita et al., 2020) requires keys to be normalized -this requirement is not met in most pretrained models. Nyströmformer (Xiong et al., 2021) and LARA (Zheng et al., 2022) do not support causal masks, which are commonly found in LLMs. Low-rank methods such as Performer (Choromanski et al., 2020) introduce substantial approximation errors, especially when they are not retrained/finetuned.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu 等NeurIPS 2024 · 被引用 479 次
- RetrievalAttention: Accelerating Long-Context LLM Inference via Vector RetrievalDi Liu, Meng Chen, Baotong Lu, Huiqiang Jiang 等NeurIPS 2025 · 被引用 148 次
- SparQ Attention: Bandwidth-Efficient LLM InferenceLuka Ribar, Ivan Chelombiev, Luke Hudlass-Galley, Charlie Blake 等ICML 2024 · 被引用 108 次
- ArkVale: Efficient Generative LLM Inference with Recallable Key-Value EvictionRenze Chen, Zhuofeng Wang, Beiquan Cao, Tong Wu 等NeurIPS 2024 · 被引用 56 次
- IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMsYuzhen Mao, Qitong Wang, Martin Ester, Ke LiICLR 2026 · 被引用 6 次
它引用的顶会 Paper10
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan 等AAAI 2021 · 被引用 675 次
相关 Paper
- SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference AccelerationJintao Zhang, Jia Wei, Pengle Zhang, Jun Zhu 等ICLR 2025
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda 等ICML 2024 · 被引用 390 次
- SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM PrefillingXiaodong Ji, Hailin Zhang, Fangcheng Fu, Bin CuiICML 2026 · 被引用 3 次
- Two Heads are Better than One: Simulating Large Transformers with Small OnesHantao Yu, Josh AlmanNeurIPS 2025 · 被引用 1 次
- ELFATT: Efficient Linear Fast Attention for Vision TransformersChong Wu, Maolin Che, Renjie Xu, Zhuoheng Ran 等ACM MM 2025 · 被引用 3 次
