On the Design of Novel Attention Mechanism for Enhanced Efficiency of Transformers
Sumit Kumar Jha, Susmit Jha, Rickard Ewetz, Alvaro Velasquez
摘要
We present a new xor-based attention function for efficient hardware implementation of transformers. While the standard attention mechanism relies on matrix multiplication between the key and the transpose of the query, we propose replacing the computation of this attention function with bitwise xor operations. We mathematically analyze the information-theoretic properties of the standard multiplication-based attention, demonstrating that it preserves input entropy, and then computationally show that the xor-based attention approximately preserves the entropy of its input despite small variations in correlations between the inputs. Across various admittedly simple tasks, including arithmetic, sorting, and text generation, we show comparable performance to baseline methods using scaled GPT models. The xor-based computation of the attention function shows substantial improvement in power consumption, latency, and circuit area compared to the corresponding multiplication-based attention function. This hardware efficiency makes xor-based attention more compelling for the deployment of transformers under tight resource constraints, opening new application domains in sustainable energy-efficient computing. Additional optimizations to the xor-based attention function can further improve efficiency of transformers.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Softermax: Hardware/Software Co-Design of an Efficient Softmax for TransformersJacob R. Stevens, Rangharajan Venkatesan, Steve Dai, Brucek Khailany 等DAC 2021 · 被引用 143 次
- FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation PredictionYubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao 等ISCA 2023 · 被引用 113 次
- ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural NetworksTae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim 等ISCA 2021 · 被引用 185 次
- EcoFormer: Energy-Saving Attention with Linear ComplexityJing Liu, Zizheng Pan, Haoyu He, Jianfei Cai 等NeurIPS 2022 · 被引用 38 次
- ELFATT: Efficient Linear Fast Attention for Vision TransformersChong Wu, Maolin Che, Renjie Xu, Zhuoheng Ran 等ACM MM 2025 · 被引用 3 次
