Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network Acceleration
Lian Liu, Zhaohui Xu, Yintao He, Ying Wang, Huawei Li, Xiaowei Li, Yinhe Han
摘要
Quantization is one of the most hardware-efficient ways to reduce inference costs for deep neural network (DNN) models. Nevertheless, with the continuous increase of DNN model sizes (240× in two years) and the emergence of large language models, existing static quantization methods fail to utilize the sparsity and redundancy of models sufficiently. Motivated by the pervasive dynamism in data tensors across DNN models, we propose a dynamic precision quantization algorithm to further reduce computational costs beyond statically quantized DNN models. Furthermore, we find that existing precision-flexible accelerators cannot support the DNN models with dynamic precision. To this end, we design a novel accelerator, Drift, and achieve online scheduling to efficiently support dynamic precision execution. We conduct experiments with various DNN models, including CNN-based and Transformer-based models. Evaluation results show that Drift achieves 2.85× speedup and 3.12× energy saving compared to existing precision-flexible accelerators with statically quantized models.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMMLian Liu, Shixin Zhao, Bing Li, Haimeng Ren 等HPCA 2025 · 被引用 15 次
- BaWA: Automatic Optimizing Pruning Metric for Large Language Models with Balanced Weight and ActivationLian Liu, Xiandong Zhao, Guanchen Li, Dong Li 等ICML 2025
- CacheEdit: Efficient Multi-round Image Editing via Adaptive Token-wise Reuse.Jinxin Yu, Xueqing Chen, Yudong Pan, Lian Liu 等ICML 2026
相关 Paper
- DRQ: Dynamic Region-based Quantization for Deep Neural Network AccelerationZhuoran Song, Bangqi Fu, Feiyang Wu, Zhaoming Jiang 等ISCA 2020 · 被引用 92 次
- INSPIRE: Accelerating Deep Neural Networks via Hardware-friendly Index-Pair EncodingFangxin Liu, Ning Yang, Zhiyan Song, Zongwu Wang 等DAC 2024 · 被引用 10 次
- Term quantization: furthering quantization at run timeHsiang-Tsung Kung, Bradley McDanel, Sai Qian ZhangSC 2020 · 被引用 10 次
- Adyna: Accelerating Dynamic Neural Networks with Adaptive SchedulingZhiyao Li, Bohan Yang, Jiaxiang Li, Taijie Chen 等HPCA 2025 · 被引用 2 次
- M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit QuantizationWeiming Hu, Zihan Zhang, Haoyan Zhang, Chen Zhang 等ASPLOS 2026 · 被引用 2 次
