Multiplication-Free Transformer Training via Piecewise Affine Operations
Atli Kosson, Martin Jaggi
Abstract
Multiplications are responsible for most of the computational cost involved in neural network training and inference. Recent research has thus looked for ways to reduce the cost associated with them. Inspired by Mogami (2020), we replace multiplication with a cheap piecewise affine approximation that is achieved by adding the bit representation of the floating point numbers together as integers. We show that transformers can be trained with the resulting modified matrix multiplications on both vision and language tasks with little to no performance impact, and without changes to the training hyperparameters. We further replace all non-linearities in the networks making them fully and jointly piecewise affine in both inputs and weights. Finally, we show that we can eliminate all multiplications in the entire training process, including operations in the forward pass, backward pass and optimizer update, demonstrating the first successful training of modern neural network architectures in a fully multiplication-free fashion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e7b036f-c3f8-4d1a-bf0d-cb445421ed3dCited by top-tier papers1
Ask how each one uses itBuilds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- Adder Attention for Vision TransformerHan Shu, Jiahao Wang, Hanting Chen, Lin Li et al.NeurIPS 2021 · 23 citations
- AdderNet: Do We Really Need Multiplications in Deep Learning?Hanting Chen, Yunhe Wang, Chunjing Xu, Boxin Shi et al.CVPR 2020
Related papers
- Is Integer Arithmetic Enough for Deep Learning Training?Alireza Ghaffari, Marzieh S. Tahaei, Mohammadreza Tayaranian, Masoud Asgharian et al.NeurIPS 2022 · 22 citations
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li et al.NeurIPS 2020 · 99 citations
- Training Transformers with 4-bit IntegersHaocheng Xi, Changhao Li, Jianfei Chen, Jun ZhuNeurIPS 2023 · 96 citations
- IM-Unpack: Training and Inference with Arbitrarily Low Precision IntegersZhanpeng Zeng, Karthikeyan Sankaralingam, Vikas SinghICML 2024 · 1 citation
- On Expressive Power of Floating-Point TransformersSejun Park, Yeachan Park, Geonho HwangICML 2026 · 2 citations
