RTN: Reparameterized Ternary Network
Yuhang Li, Xin Dong, Sai Qian Zhang, Haoli Bai, Yuanpeng Chen, Wei Wang
Abstract
To deploy deep neural networks on resource-limited devices, quantization has been widely explored. In this work, we study the extremely low-bit networks which have tremendous speed-up, memory saving with quantized activation and weights. We first bring up three omitted issues in extremely low-bit networks: the squashing range of quantized values; the gradient vanishing during backpropagation and the unexploited hardware acceleration of ternary networks. By reparameterizing quantized activation and weights vector with full precision scale and offset for fixed ternary vector, we decouple the range and magnitude from the direction to extenuate the three issues. Learnable scale and offset can automatically adjust the range of quantized values and sparsity without gradient vanishing. A novel encoding and computation pattern are designed to support efficient computing for our reparameterized ternary network (RTN). Experiments on ResNet-18 for ImageNet demonstrate that the proposed RTN finds a much better efficiency between bitwidth and accuracy, and achieves up to 26.76% relative accuracy improvement compared with state-of-the-art methods. Moreover, we validate the proposed computation pattern on Field Programmable Gate Arrays (FPGA), and it brings 46.46× and 89.17× savings on power and area respectively compared with the full precision convolution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad5819f3-2715-4dde-9604-dccf567a2b75Cited by top-tier papers6
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 66 citations
- Towards Efficient Post-training Quantization of Pre-trained Language ModelsHaoli Bai, Lu Hou, Lifeng Shang, Xin Jiang et al.NeurIPS 2022 · 62 citations
- TRQ: Ternary Neural Networks With Residual QuantizationYue Li, Wenrui Ding, Chunlei Liu, Baochang Zhang et al.AAAI 2021 · 34 citations
- FATNN: Fast and Accurate Ternary Neural Networks*Peng Chen, Bohan Zhuang, Chunhua ShenICCV 2021 · 22 citations
- Count2Multiply: Reliable In-Memory High-Radix CountingJoão Paulo C. de Lima, Benjamin F. Morris III, Asif Ali Khan, Jerónimo Castrillón et al.HPCA 2026
Builds on1
Related papers
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksYikai Wang, Yi Yang, Fuchun Sun, Anbang YaoICCV 2021 · 18 citations
- Harmonious Coexistence of Structured Weight Pruning and Ternarization for Deep Neural NetworksLi Yang, Zhezhi He, Deliang FanAAAI 2020 · 28 citations
- Learnable Companding Quantization for Accurate Low-Bit Neural NetworksKohei YamamotoCVPR 2021
- Improving Low-Precision Network Quantization via Bin RegularizationTiantian Han, Dong Li, Ji Liu, Lu Tian et al.ICCV 2021 · 45 citations
- S: Sign-Sparse-Shift Reparametrization for Effective Training of Low-bit Shift NetworksXinlin Li, Bang Liu, Yaoliang Yu, Wulong Liu et al.NeurIPS 2021 · 12 citations
