WrapNet: Neural Net Inference with Ultra-Low-Precision Arithmetic
Renkun Ni, Hong-Min Chu, Oscar Castañeda, Ping-yeh Chiang, Christoph Studer, Tom Goldstein
Abstract
Low-resolution neural networks represent both weights and activations with few bits, drastically reducing the multiplication complexity. Nonetheless, these products are accumulated using high-resolution (typically 32-bit) additions, an operation that dominates the arithmetic complexity of inference when using extreme quantization (e.g., binary weights). To further optimize inference, we propose a method that adapts neural networks to use low-resolution (8-bit) additions in the accumulators, achieving classification accuracy comparable to their 32-bit counterparts. We achieve resilience to low-resolution accumulation by inserting a cyclic activation layer, as well as an overflow penalty regularizer. We demonstrate the efficacy of our approach on both software and hardware platforms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0cb69b67-5dda-40cd-b50d-86f9bc072422Cited by top-tier papers2
- A2Q+: Improving Accumulator-Aware Weight QuantizationIan Colbert, Alessandro Pappalardo, Jakoba Petri-Koenig, Yaman UmurogluICML 2024 · 11 citations
- Understanding Neural Network Binarization with Forward and Backward Proximal QuantizersYiwei Lu, Yaoliang Yu, Xinlin Li, Vahid Partovi NiaNeurIPS 2023 · 5 citations
Builds on1
Related papers
- Towards Cheaper Inference in Deep Networks with Lower Bit-Width AccumulatorsYaniv Blumenfeld, Itay Hubara, Daniel SoudryICLR 2024 · 5 citations
- A2Q: Accumulator-Aware Quantization with Guaranteed Overflow AvoidanceIan Colbert, Alessandro Pappalardo, Jakoba Petri-KoenigICCV 2023 · 19 citations
- Linear Symmetric Quantization of Neural Networks for Low-precision Integer HardwareXiandong Zhao, Ying Wang, Xuyi Cai, Cheng Liu et al.ICLR 2020 · 67 citations
- Toward Efficient Low-Precision Training: Data Format Optimization and Hysteresis QuantizationSunwoo Lee, Jeongwoo Park, Dongsuk JeonICLR 2022 · 11 citations
- RTN: Reparameterized Ternary NetworkYuhang Li, Xin Dong, Sai Qian Zhang, Haoli Bai et al.AAAI 2020 · 34 citations
