FATNN: Fast and Accurate Ternary Neural Networks*
Peng Chen, Bohan Zhuang, Chunhua Shen
摘要
Ternary Neural Networks (TNNs) have received much attention due to being potentially orders of magnitude faster in inference, as well as more power efficient, than full-precision counterparts. However, 2 bits are required to encode the ternary representation with only 3 quantization levels leveraged. As a result, conventional TNNs have similar memory consumption and speed compared with the standard 2-bit models, but have worse representational capability. Moreover, there is still a significant gap in accuracy between TNNs and full-precision networks, hampering their deployment to real applications. To tackle these two challenges, in this work, we first show that, under some mild constraints, computational complexity of the ternary inner product can be reduced by 2×. Second, to mitigate the performance gap, we elaborately design an implementation-dependent ternary quantization algorithm. The proposed framework is termed Fast and Accurate Ternary Neural Networks (FATNN). Experiments on image classification demonstrate that our FATNN surpasses the state-of-the-arts by a significant margin in accuracy. More importantly, speedup evaluation compared with various precision is analyzed on several platforms, which serves as a strong benchmark for further research. Source code and models are available at: https://github.com/MonashAI/QTool
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- RTN: Reparameterized Ternary NetworkYuhang Li, Xin Dong, Sai Qian Zhang, Haoli Bai 等AAAI 2020 · 被引用 34 次
- Forward and Backward Information Retention for Accurate Binary Neural NetworksHaotong Qin, Ruihao Gong, Xianglong Liu, Mingzhu Shen 等CVPR 2020
相关 Paper
- TRQ: Ternary Neural Networks With Residual QuantizationYue Li, Wenrui Ding, Chunlei Liu, Baochang Zhang 等AAAI 2021 · 被引用 34 次
- ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM QuantizationZechun Liu, Changsheng Zhao, Hanxian Huang, Sijia Chen 等NeurIPS 2025 · 被引用 50 次
- CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMsShigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan 等ICML 2026 · 被引用 2 次
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksYikai Wang, Yi Yang, Fuchun Sun, Anbang YaoICCV 2021 · 被引用 18 次
- FAST: DNN Training Under Variable Precision Block Floating Point with Stochastic RoundingSai Qian Zhang, Bradley McDanel, H. T. KungHPCA 2022 · 被引用 68 次
