CTA: Hardware-Software Co-design for Compressed Token Attention Mechanism
Haoran Wang, Haobo Xu, Ying Wang, Yinhe Han
Abstract
The attention mechanism is becoming an integral part of modern neural networks, bringing breakthroughs to Natural Language Processing (NLP) applications and even Computer Vision (CV) applications. Unfortunately, the superiority of attention mechanism comes from its ability to model relations between any two positions in long sequence, which incurs high inference overhead. For state-of-the-art AI workloads such as Bert or GPT-2, attention mechanism is reported to account up to 50% of the inference overhead. Previous works seek to alleviate this performance bottleneck by removing useless relations for each position and accelerate position-specific operations. However their attempts require selecting from a sequence of relations once for each position, which is essentially frequent on-the-fly pruning and breaks the inherent parallelism in attention mechanism. In this paper, we propose CTA, an algorithm-architecture co-designed solution that can substantially reduce theoretic complexity of attention mechanism, enabling significant speedup and energy saving. Inspired by the fact that the feature sequence encoded by attention mechanism contain a large number of semantic feature repetition, we propose a novel approximation scheme that can efficiently remove that repetition, only calculating attention among necessary features thus reducing computation complexity quadratically. To utilize this algorithmic bonus and empower high performance attention mechanism inference, we devise specialized architecture to efficiently support the proposed approximation scheme. Extensive experiments show that, on average, CTA achieves 27.7× speedup, 634.0× energy savings with no accuracy loss, and 44.2× speedup, 950.0× energy savings with around 1% accuracy loss over Nvidia V100-SXM2 GPU. Also, CTA achieves 22.8× speedup, 479.6× energy savings over ELSA accelerator+GPU system.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6cec2ca6-8566-4c94-9053-ff0a43d7cc0dCited by top-tier papers4
- PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing SystemYintao He, Haiyu Mao, Christina Giannoula, Mohammad Sadrosadati et al.ASPLOS 2025 · 37 citations
- TB-STC: Transposable Block-wise N: M Structured Sparse Tensor CoreJun Liu, Shulin Zeng, Junbo Zhao, Li Ding et al.HPCA 2025 · 9 citations
- PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage FusionHuizheng Wang, Hongbin Wang, Zichuan Wang, Zhiheng Yue et al.HPCA 2026 · 2 citations
- Accelerating Sparse Transformer Inference on GPUWenhao Dai, Haodong Deng, Mengfei Rong, Xinyu Yang et al.PPoPP 2026 · 1 citation
Related papers
- ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural NetworksTae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim et al.ISCA 2021 · 185 citations
- Blaze: An Efficient Bit-Sparse Attention Architecture With Workload Orchestration OptimizationRunzhou Zhang, Faxian Sun, Yiming Wang, Kunchen Zou et al.DAC 2025
- DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured PruningXiao Xiong, Zhaorui Chen, Yue Liang, Minghao Tian et al.ASPLOS 2025 · 3 citations
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 412 citations
- DOTA: detect and omit weak attentions for scalable transformer accelerationZheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen et al.ASPLOS 2022 · 131 citations
