Latency-Constrained Input-Aware Quantization of Time Series Inference Workflows at the Edge
Ashitabh Misra, Nurani Saoda, Tarek F. Abdelzaher
摘要
Quantization is a common technique for compressing Deep Neural Networks (DNNs) for edge device deployment, crucial for real-time time-series sensor data classification. While Adaptive Quantization (AQ) techniques allow updating the model's quantization scheme based on resource constraints without re-training, they are input-agnostic. Conversely, Input-aware Quantization (IQ) techniques adjust quantization per input, improving accuracy but ignoring resource constraints. Additionally, most quantization methods are designed for vision-based applications and underutilize time-frequency domain semantics. To this end, we present ReactQuant, an end-to-end framework for input-aware quantization of DNNs with strict resource constraints, specialized for time-series applications. ReactQuant features Partition-Derived Quantization, which segments time-series input based on its time and frequency domain features; a two-staged training process that restricts quantization space to each layer's promising bit-width candidates; and a novel algorithm to perform input-aware quantization within user-defined resource constraints. Extensive experimentation shows that ReactQuant outperforms AQ and IQ techniques in accuracy and ensures strict compliance with user-defined resource constraints. Under similar resource constraints, ReactQuant achieves up to 19% higher accuracy than AQ techniques and 4.6% higher than state-of-the-art IQ technique. It incurs no resource constraint violations, compared to up to 27.63% for state-of-the-art IQ technique.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- QEBVerif: Quantization Error Bound Verification of Neural NetworksYedi Zhang, Fu Song, Jun SunCAV 2023 · 被引用 17 次
- Adaptive Loss-Aware Quantization for Multi-Bit NetworksZhongnan Qu, Zimu Zhou, Yun Cheng, Lothar ThieleCVPR 2020
- SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian ApproximationCong Guo, Yuxian Qiu, Jingwen Leng, Xiaotian Gao 等ICLR 2022 · 被引用 92 次
- Fixed-Point Back-Propagation TrainingXishan Zhang, Shaoli Liu, Rui Zhang, Chang Liu 等CVPR 2020
- Data-Free Network Compression via Parametric Non-uniform Mixed Precision QuantizationVladimir Chikin, Mikhail AntiukhCVPR 2022 · 被引用 17 次
