SC2020Top-tier venue
Term quantization: furthering quantization at run time
Hsiang-Tsung Kung, Bradley McDanel, Sai Qian Zhang
Abstract
We present a novel technique, called Term Quantization (TQ), for furthering quantization at run time for improved computational efficiency of deep neural networks (DNNs) already quantized with conventional quantization methods. TQ operates on power-of-two terms in expressions of values. In computing a dot-product computation, TQ dynamically selects a fixed number of largest terms to use from values of the two vectors. By exploiting weight and data distributions typically present in DNNs, TQ has a minimal impact on DNN model performance (e.g., accuracy or perplexity). We use TQ to facilitate tightly synchronized processor arrays, such as systolic arrays, for efficient parallel processing. We evaluate TQ on an MLP for MNIST, multiple CNNs for ImageNet and an LSTM for Wikitext-2. We demonstrate significant reductions in inference computation costs (between 3-10×) compared to conventional uniform quantization for the same level of model performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f81330a5-ccbb-4159-a166-af55a2f02c24Cited by top-tier papers2
- PICACHU: Plug-In CGRA Handling Upcoming Nonlinear Operations in LLMsJiajun Qin, Tianhua Xia, Cheng Tan, Jeff Zhang et al.ASPLOS 2025 · 17 citations
- MaverIQ: Fingerprint-Guided Extrapolation and Fragmentation-Aware Layering for Intent-Based LLM ServingDimitrios Liakopoulos, Prasoon Sinha, Tianrui Hu, Myungjin Lee et al.SC 2025 · 2 citations
Related papers
- Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network AccelerationLian Liu, Zhaohui Xu, Yintao He, Ying Wang et al.DAC 2024 · 5 citations
- BiQGEMM: matrix multiplication with lookup table for binary-coding-based quantized DNNsYongkweon Jeon, Baeseong Park, Se Jung Kwon, Byeongwook Kim et al.SC 2020 · 31 citations
- INSPIRE: Accelerating Deep Neural Networks via Hardware-friendly Index-Pair EncodingFangxin Liu, Ning Yang, Zhiyan Song, Zongwu Wang et al.DAC 2024 · 10 citations
- Training for multi-resolution inference using reusable quantization termsSai Qian Zhang, Bradley McDanel, H. T. Kung, Xin DongASPLOS 2021 · 9 citations
- Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision ActivationsYichi Zhang, Ritchie Zhao, Weizhe Hua, Nayun Xu et al.ICLR 2020 · 28 citations
