TRQ: Ternary Neural Networks With Residual Quantization
Yue Li, Wenrui Ding, Chunlei Liu, Baochang Zhang, Guodong Guo
Abstract
Ternary neural networks (TNNs) are potential for network acceleration by reducing the full-precision weights in network to ternary ones, e.g., -1, 0, 1. However, existing TNNs are mostly calculated based on rule-of-thumb quantization methods by simply thresholding operations, which causes a significant accuracy loss. In this paper, we introduce a stemresidual framework which provides new insight into ternary quantization, termed Ternary Residual Quantization (TRQ), to achieve more powerful TNNs. Rather than directly thresholding operations, TRQ recursively performs quantization on full-precision weights for a refined reconstruction by combining the binarized stem and residual parts.With such a unique quantization process, TRQ endows the quantizer with high flexibility and precision. Furthermore, our TRQ is generic, which can be easily extended to multiple bits through recursively encoded residual for a better recognition accuracy. Extensive experimental results demonstrate that the proposed method yields great recognition accuracy while being accelerated.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64d1efe0-d9fa-4919-bffb-27cfe4a09531Cited by top-tier papers4
- Autoregressive Image Generation using Residual QuantizationDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho et al.CVPR 2022 · 184 citations
- One-Step Forward and Backtrack: Overcoming Zig-Zagging in Loss-Aware Quantization TrainingLianbo Ma, Yuee Zhou, Jianlun Ma, Guo Yu et al.AAAI 2024 · 5 citations
- Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative LearningQihua Zhou, Song Guo, Yi Liu, Jie Zhang et al.NeurIPS 2022 · 4 citations
- MoMask: Generative Masked Modeling of 3D Human MotionsChuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang et al.CVPR 2024
Builds on6
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Searching for Low-Bit Weights in Quantized Neural NetworksZhaohui Yang, Yunhe Wang, Kai Han, Chunjing Xu et al.NeurIPS 2020 · 103 citations
- Training Binary Neural Networks through Learning with Noisy SupervisionKai Han, Yunhe Wang, Yixing Xu, Chunjing Xu et al.ICML 2020 · 63 citations
- Bayesian Optimized 1-Bit CNNsJiaxin Gu, Junhe Zhao, Xiaolong Jiang, Baochang Zhang et al.ICCV 2019 · 57 citations
- BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary ActivationsHyungjun Kim, Kyungsu Kim, Jinseok Kim, Jae-Joon KimICLR 2020 · 51 citations
Related papers
- FATNN: Fast and Accurate Ternary Neural Networks*Peng Chen, Bohan Zhuang, Chunhua ShenICCV 2021 · 22 citations
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksYikai Wang, Yi Yang, Fuchun Sun, Anbang YaoICCV 2021 · 18 citations
- RTN: Reparameterized Ternary NetworkYuhang Li, Xin Dong, Sai Qian Zhang, Haoli Bai et al.AAAI 2020 · 34 citations
- Harmonious Coexistence of Structured Weight Pruning and Ternarization for Deep Neural NetworksLi Yang, Zhezhi He, Deliang FanAAAI 2020 · 28 citations
- APNN-TC: accelerating arbitrary precision neural networks on ampere GPU tensor coresBoyuan Feng, Yuke Wang, Tong Geng, Ang Li et al.SC 2021 · 36 citations
