TRQ: Ternary Neural Networks With Residual Quantization
Yue Li, Wenrui Ding, Chunlei Liu, Baochang Zhang, Guodong Guo
摘要
Ternary neural networks (TNNs) are potential for network acceleration by reducing the full-precision weights in network to ternary ones, e.g., -1, 0, 1. However, existing TNNs are mostly calculated based on rule-of-thumb quantization methods by simply thresholding operations, which causes a significant accuracy loss. In this paper, we introduce a stemresidual framework which provides new insight into ternary quantization, termed Ternary Residual Quantization (TRQ), to achieve more powerful TNNs. Rather than directly thresholding operations, TRQ recursively performs quantization on full-precision weights for a refined reconstruction by combining the binarized stem and residual parts.With such a unique quantization process, TRQ endows the quantizer with high flexibility and precision. Furthermore, our TRQ is generic, which can be easily extended to multiple bits through recursively encoded residual for a better recognition accuracy. Extensive experimental results demonstrate that the proposed method yields great recognition accuracy while being accelerated.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Autoregressive Image Generation using Residual QuantizationDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho 等CVPR 2022 · 被引用 184 次
- One-Step Forward and Backtrack: Overcoming Zig-Zagging in Loss-Aware Quantization TrainingLianbo Ma, Yuee Zhou, Jianlun Ma, Guo Yu 等AAAI 2024 · 被引用 5 次
- Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative LearningQihua Zhou, Song Guo, Yi Liu, Jie Zhang 等NeurIPS 2022 · 被引用 4 次
- MoMask: Generative Masked Modeling of 3D Human MotionsChuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang 等CVPR 2024
它引用的顶会 Paper6
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- Searching for Low-Bit Weights in Quantized Neural NetworksZhaohui Yang, Yunhe Wang, Kai Han, Chunjing Xu 等NeurIPS 2020 · 被引用 103 次
- Training Binary Neural Networks through Learning with Noisy SupervisionKai Han, Yunhe Wang, Yixing Xu, Chunjing Xu 等ICML 2020 · 被引用 63 次
- Bayesian Optimized 1-Bit CNNsJiaxin Gu, Junhe Zhao, Xiaolong Jiang, Baochang Zhang 等ICCV 2019 · 被引用 57 次
- BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary ActivationsHyungjun Kim, Kyungsu Kim, Jinseok Kim, Jae-Joon KimICLR 2020 · 被引用 51 次
相关 Paper
- FATNN: Fast and Accurate Ternary Neural Networks*Peng Chen, Bohan Zhuang, Chunhua ShenICCV 2021 · 被引用 22 次
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksYikai Wang, Yi Yang, Fuchun Sun, Anbang YaoICCV 2021 · 被引用 18 次
- RTN: Reparameterized Ternary NetworkYuhang Li, Xin Dong, Sai Qian Zhang, Haoli Bai 等AAAI 2020 · 被引用 34 次
- Harmonious Coexistence of Structured Weight Pruning and Ternarization for Deep Neural NetworksLi Yang, Zhezhi He, Deliang FanAAAI 2020 · 被引用 28 次
- APNN-TC: accelerating arbitrary precision neural networks on ampere GPU tensor coresBoyuan Feng, Yuke Wang, Tong Geng, Ang Li 等SC 2021 · 被引用 36 次
