ClusterFormer: Neural Clustering Attention for Efficient and Effective Transformer
Ningning Wang, Guobing Gan, Peng Zhang, Shuai Zhang, Victor Junqiu Wei, Qun Liu, Xin Jiang
Abstract
Recently, a lot of research has been carried out to improve the efficiency of Transformer. Among them, the sparse pattern-based method is an important branch of efficient Transformers. However, some existing sparse methods usually use fixed patterns to select words, without considering similarities between words. Other sparse methods use clustering patterns to select words, but the clustering process is separate from the training process of the target task, which causes a decrease in effectiveness. To address these limitations, we design a neural clustering method, which can be seamlessly integrated into the Self-Attention Mechanism in Transformer. The clustering task and the target task are jointly trained and optimized to benefit each other, leading to significant effectiveness improvement. In addition, our method groups the words with strong dependencies into the same cluster and performs the attention mechanism for each cluster independently, which improves the efficiency. We verified our method on machine translation, text classification, natural language inference, and text matching tasks. Experimental results show that our method outperforms two typical sparse attention methods, Reformer and Routing Transformer while having a comparable or even better time and memory efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 858fd6d5-57ba-4c39-bf5d-74c70f1bab1cCited by top-tier papers3
- Transformer-VQ: Linear-Time Transformers via Vector QuantizationLucas D. LingleICLR 2024 · 30 citations
- DARKER: Efficient Transformer with Data-driven Attention Mechanism for Time SeriesRundong Zuo, Guozhong Li, Rui Cao, Byron Choi et al.VLDB 2024 · 4 citations
- Treeformer: Dense Gradient Trees for Efficient Attention ComputationLovish Madaan, Srinadh Bhojanapalli, Himanshu Jain, Prateek JainICLR 2023
Builds on4
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
- Multiple Positional Self-Attention Network for Text ClassificationBiyun Dai, Jinlong Li, Ruoyi XuAAAI 2020 · 10 citations
Related papers
- SpARC: Token Similarity-Aware Sparse Attention Transformer Accelerator via Row-wise ClusteringHan Cho, Dongjun Kim, Seung-Eon Hwang, Jongsun ParkDAC 2024 · 7 citations
- ClusterAttn: KV Cache Compression under Intrinsic Attention ClusteringMinwei Zhang, Haifeng Sun, Jingyu Wang, Shaolong Li et al.ACL 2025 · 5 citations
- Neural Machine Translation with Joint RepresentationYanyang Li, Qiang Wang, Tong Xiao, Tongran Liu et al.AAAI 2020 · 10 citations
- Diffuser: Efficient Transformers with Multi-Hop Attention Diffusion for Long SequencesAosong Feng, Irene Li, Yuang Jiang, Rex YingAAAI 2023 · 20 citations
- ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural NetworksTae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim et al.ISCA 2021 · 185 citations
