RTN: Reparameterized Ternary Network
Yuhang Li, Xin Dong, Sai Qian Zhang, Haoli Bai, Yuanpeng Chen, Wei Wang
摘要
To deploy deep neural networks on resource-limited devices, quantization has been widely explored. In this work, we study the extremely low-bit networks which have tremendous speed-up, memory saving with quantized activation and weights. We first bring up three omitted issues in extremely low-bit networks: the squashing range of quantized values; the gradient vanishing during backpropagation and the unexploited hardware acceleration of ternary networks. By reparameterizing quantized activation and weights vector with full precision scale and offset for fixed ternary vector, we decouple the range and magnitude from the direction to extenuate the three issues. Learnable scale and offset can automatically adjust the range of quantized values and sparsity without gradient vanishing. A novel encoding and computation pattern are designed to support efficient computing for our reparameterized ternary network (RTN). Experiments on ResNet-18 for ImageNet demonstrate that the proposed RTN finds a much better efficiency between bitwidth and accuracy, and achieves up to 26.76% relative accuracy improvement compared with state-of-the-art methods. Moreover, we validate the proposed computation pattern on Field Programmable Gate Arrays (FPGA), and it brings 46.46× and 89.17× savings on power and area respectively compared with the full precision convolution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 被引用 66 次
- Towards Efficient Post-training Quantization of Pre-trained Language ModelsHaoli Bai, Lu Hou, Lifeng Shang, Xin Jiang 等NeurIPS 2022 · 被引用 62 次
- TRQ: Ternary Neural Networks With Residual QuantizationYue Li, Wenrui Ding, Chunlei Liu, Baochang Zhang 等AAAI 2021 · 被引用 34 次
- FATNN: Fast and Accurate Ternary Neural Networks*Peng Chen, Bohan Zhuang, Chunhua ShenICCV 2021 · 被引用 22 次
- Count2Multiply: Reliable In-Memory High-Radix CountingJoão Paulo C. de Lima, Benjamin F. Morris III, Asif Ali Khan, Jerónimo Castrillón 等HPCA 2026
它引用的顶会 Paper1
相关 Paper
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksYikai Wang, Yi Yang, Fuchun Sun, Anbang YaoICCV 2021 · 被引用 18 次
- Harmonious Coexistence of Structured Weight Pruning and Ternarization for Deep Neural NetworksLi Yang, Zhezhi He, Deliang FanAAAI 2020 · 被引用 28 次
- Learnable Companding Quantization for Accurate Low-Bit Neural NetworksKohei YamamotoCVPR 2021
- Improving Low-Precision Network Quantization via Bin RegularizationTiantian Han, Dong Li, Ji Liu, Lu Tian 等ICCV 2021 · 被引用 45 次
- S: Sign-Sparse-Shift Reparametrization for Effective Training of Low-bit Shift NetworksXinlin Li, Bang Liu, Yaoliang Yu, Wulong Liu 等NeurIPS 2021 · 被引用 12 次
