IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs
Yuzhen Mao, Martin Ester, Ke Li
Abstract
One limitation of existing Transformer-based models is that they cannot handle very long sequences as input since their self-attention operations exhibit quadratic time and space complexity. This problem becomes especially acute when Transformers are deployed on hardware platforms equipped only with CPUs. To address this issue, we propose a novel method for accelerating self-attention at inference time that works with pretrained Transformer models out-of-the-box without requiring retraining. We experiment using our method to accelerate various long-sequence Transformers, including a leading LLaMA 2-based LLM, on various benchmarks and demonstrate a speedup of 2.73 × -7.63× while retaining 98.6% -99.6% of the accuracy of the original pretrained models. The code is available on our project website at https://yuzhenmao.github.io/IceFormer/ . INTRODUCTION Transformers (Vaswani et al., 2017) have powered incredible advances in NLP, as exemplified by large language models (LLMs) such as GPT-4 and LLaMA 2. Increasingly LLMs are applied to exceptionally long input sequences, which enables many exciting applications such as long-form content creation, extended conversations, and large document search and analysis (OpenAI, 2023; Anthropic, 2023). While LLMs can be feasibly trained with expensive hardware accelerators (e.g. GPUs), they need to be deployed on commodity devices, which may only be equipped with CPUs. However, it is currently challenging to deploy LLMs on CPUs due to their high computation cost (Dice & Kogan, 2021) . A significant computational bottleneck arises from the self-attention mechanism that is integral to Transformers -both time and space complexity are quadratic in the sequence length. This problem is exacerbated in the context of LLMs, which are often used on very long sequences. To handle long input sequences, there has been substantial research into reducing the quadratic time complexity of self-attention -these methods are collectively known as efficient Transformers. However, many do not meet the needs of LLMs and are therefore difficult to apply to LLMs. An ideal acceleration method for LLMs should satisfy four criteria: (1) No retraining -the method should not require the model to be retrained, given the enormous computational expense of training LLMs; (2) Generality -the method should be applicable to a variety of LLMs, rather than just those trained with particular constraints built-in; (3) High accuracy -the method should not introduce large approximation errors, since LLMs have many attention layers and so errors from earlier layers can compound; (4) Fast inference -the method should achieve fast test-time performance. Satisfying all these criteria simultaneously is difficult, and to our knowledge no existing methods can do so. For example, Transformers with fixed attention patterns, e.g., Longformer (Beltagy et al., 2020) , require retraining the model before they can be used. Reformer (Nikita et al., 2020) requires keys to be normalized -this requirement is not met in most pretrained models. Nyströmformer (Xiong et al., 2021) and LARA (Zheng et al., 2022) do not support causal masks, which are commonly found in LLMs. Low-rank methods such as Performer (Choromanski et al., 2020) introduce substantial approximation errors, especially when they are not retrained/finetuned.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0559fdad-9986-4dd5-87a4-199ae21102d8Cited by top-tier papers7
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu et al.NeurIPS 2024 · 479 citations
- RetrievalAttention: Accelerating Long-Context LLM Inference via Vector RetrievalDi Liu, Meng Chen, Baotong Lu, Huiqiang Jiang et al.NeurIPS 2025 · 148 citations
- SparQ Attention: Bandwidth-Efficient LLM InferenceLuka Ribar, Ivan Chelombiev, Luke Hudlass-Galley, Charlie Blake et al.ICML 2024 · 108 citations
- ArkVale: Efficient Generative LLM Inference with Recallable Key-Value EvictionRenze Chen, Zhuofeng Wang, Beiquan Cao, Tong Wu et al.NeurIPS 2024 · 56 citations
- IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMsYuzhen Mao, Qitong Wang, Martin Ester, Ke LiICLR 2026 · 6 citations
Builds on10
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
Related papers
- SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference AccelerationJintao Zhang, Jia Wei, Pengle Zhang, Jun Zhu et al.ICLR 2025
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda et al.ICML 2024 · 390 citations
- SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM PrefillingXiaodong Ji, Hailin Zhang, Fangcheng Fu, Bin CuiICML 2026 · 3 citations
- Two Heads are Better than One: Simulating Large Transformers with Small OnesHantao Yu, Josh AlmanNeurIPS 2025 · 1 citation
- ELFATT: Efficient Linear Fast Attention for Vision TransformersChong Wu, Maolin Che, Renjie Xu, Zhuoheng Ran et al.ACM MM 2025 · 3 citations
