Is Integer Arithmetic Enough for Deep Learning Training?
Alireza Ghaffari, Marzieh S. Tahaei, Mohammadreza Tayaranian, Masoud Asgharian, Vahid Partovi Nia
摘要
The ever-increasing computational complexity of deep learning models makes their training and deployment difficult on various cloud and edge platforms. Replacing floating-point arithmetic with low-bit integer arithmetic is a promising approach to save energy, memory footprint, and latency of deep learning models. As such, quantization has attracted the attention of researchers in recent years. However, using integer numbers to form a fully functional integer training pipeline including forward pass, back-propagation, and stochastic gradient descent is not studied in detail. Our empirical and mathematical results reveal that integer arithmetic seems to be enough to train deep learning models. Unlike recent proposals, instead of quantization, we directly switch the number representation of computations. Our novel training method forms a fully integer training pipeline that does not change the trajectory of the loss and accuracy compared to floating-point, nor does it need any special hyper-parameter tuning, distribution adjustment, or gradient clipping. Our experimental results show that our proposed method is effective in a wide variety of tasks such as classification (including vision transformers), object detection, and semantic segmentation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Understanding Neural Network Binarization with Forward and Backward Proximal QuantizersYiwei Lu, Yaoliang Yu, Xinlin Li, Vahid Partovi NiaNeurIPS 2023 · 被引用 5 次
- Unforgeability in Stochastic Gradient DescentTeodora Baluta, Ivica Nikolic, Racchit Jain, Divesh Aggarwal 等CCS 2023
- HOT: Hadamard-based Optimized TrainingSeonggon Kim, Juncheol Shin, Seung-taek Woo, Eunhyeok ParkCVPR 2025
它引用的顶会 Paper6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating PointBita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Ming Liu 等NeurIPS 2020 · 被引用 153 次
- Distribution Adaptive INT8 Quantization for Training CNNsKang Zhao, Sida Huang, Pan Pan, Yinghan Li 等AAAI 2021 · 被引用 86 次
- F8Net: Fixed-Point 8-bit Only Multiplication for Network QuantizationQing Jin, Jian Ren, Richard Zhuang, Sumant Hanumante 等ICLR 2022 · 被引用 57 次
- Fixed-Point Back-Propagation TrainingXishan Zhang, Shaoli Liu, Rui Zhang, Chang Liu 等CVPR 2020
相关 Paper
- Winning Both the Accuracy of Floating Point Activation and the Simplicity of Integer ArithmeticYulhwa Kim, Jaeyong Jang, Jehun Lee, Jihoon Park 等ICLR 2023
- Shifted and Squeezed 8-bit Floating Point format for Low-Precision Training of Deep Neural NetworksLéopold Cambier, Anahita Bhiwandiwalla, Ting Gong, Oguz H. Elibol 等ICLR 2020 · 被引用 53 次
- 8-bit Transformer Inference and Fine-tuning for Edge AcceleratorsJeffrey Yu, Kartik Prabhu, Yonatan Urman, Robert M. Radway 等ASPLOS 2024 · 被引用 26 次
- I-BERT: Integer-only BERT QuantizationSehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney 等ICML 2021 · 被引用 439 次
- BOLD: Boolean Logic Deep LearningVan Minh Nguyen, Cristian Ocampo-Blandon, Aymen Askri, Louis Leconte 等NeurIPS 2024 · 被引用 4 次
