FATNN: Fast and Accurate Ternary Neural Networks*
Peng Chen, Bohan Zhuang, Chunhua Shen
Abstract
Ternary Neural Networks (TNNs) have received much attention due to being potentially orders of magnitude faster in inference, as well as more power efficient, than full-precision counterparts. However, 2 bits are required to encode the ternary representation with only 3 quantization levels leveraged. As a result, conventional TNNs have similar memory consumption and speed compared with the standard 2-bit models, but have worse representational capability. Moreover, there is still a significant gap in accuracy between TNNs and full-precision networks, hampering their deployment to real applications. To tackle these two challenges, in this work, we first show that, under some mild constraints, computational complexity of the ternary inner product can be reduced by 2×. Second, to mitigate the performance gap, we elaborately design an implementation-dependent ternary quantization algorithm. The proposed framework is termed Fast and Accurate Ternary Neural Networks (FATNN). Experiments on image classification demonstrate that our FATNN surpasses the state-of-the-arts by a significant margin in accuracy. More importantly, speedup evaluation compared with various precision is analyzed on several platforms, which serves as a strong benchmark for further research. Source code and models are available at: https://github.com/MonashAI/QTool
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b00a914e-9e4e-4252-a4d3-14b4d9a10276Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- RTN: Reparameterized Ternary NetworkYuhang Li, Xin Dong, Sai Qian Zhang, Haoli Bai et al.AAAI 2020 · 34 citations
- Forward and Backward Information Retention for Accurate Binary Neural NetworksHaotong Qin, Ruihao Gong, Xianglong Liu, Mingzhu Shen et al.CVPR 2020
Related papers
- TRQ: Ternary Neural Networks With Residual QuantizationYue Li, Wenrui Ding, Chunlei Liu, Baochang Zhang et al.AAAI 2021 · 34 citations
- ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM QuantizationZechun Liu, Changsheng Zhao, Hanxian Huang, Sijia Chen et al.NeurIPS 2025 · 50 citations
- CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMsShigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan et al.ICML 2026 · 2 citations
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksYikai Wang, Yi Yang, Fuchun Sun, Anbang YaoICCV 2021 · 18 citations
- FAST: DNN Training Under Variable Precision Block Floating Point with Stochastic RoundingSai Qian Zhang, Bradley McDanel, H. T. KungHPCA 2022 · 68 citations
