Latency-Constrained Input-Aware Quantization of Time Series Inference Workflows at the Edge
Ashitabh Misra, Nurani Saoda, Tarek F. Abdelzaher
Abstract
Quantization is a common technique for compressing Deep Neural Networks (DNNs) for edge device deployment, crucial for real-time time-series sensor data classification. While Adaptive Quantization (AQ) techniques allow updating the model's quantization scheme based on resource constraints without re-training, they are input-agnostic. Conversely, Input-aware Quantization (IQ) techniques adjust quantization per input, improving accuracy but ignoring resource constraints. Additionally, most quantization methods are designed for vision-based applications and underutilize time-frequency domain semantics. To this end, we present ReactQuant, an end-to-end framework for input-aware quantization of DNNs with strict resource constraints, specialized for time-series applications. ReactQuant features Partition-Derived Quantization, which segments time-series input based on its time and frequency domain features; a two-staged training process that restricts quantization space to each layer's promising bit-width candidates; and a novel algorithm to perform input-aware quantization within user-defined resource constraints. Extensive experimentation shows that ReactQuant outperforms AQ and IQ techniques in accuracy and ensures strict compliance with user-defined resource constraints. Under similar resource constraints, ReactQuant achieves up to 19% higher accuracy than AQ techniques and 4.6% higher than state-of-the-art IQ technique. It incurs no resource constraint violations, compared to up to 27.63% for state-of-the-art IQ technique.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 09244677-461b-4fa5-a358-c4f94b5d57a8Cited by top-tier papers1
Ask how each one uses itRelated papers
- QEBVerif: Quantization Error Bound Verification of Neural NetworksYedi Zhang, Fu Song, Jun SunCAV 2023 · 17 citations
- Adaptive Loss-Aware Quantization for Multi-Bit NetworksZhongnan Qu, Zimu Zhou, Yun Cheng, Lothar ThieleCVPR 2020
- SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian ApproximationCong Guo, Yuxian Qiu, Jingwen Leng, Xiaotian Gao et al.ICLR 2022 · 92 citations
- Fixed-Point Back-Propagation TrainingXishan Zhang, Shaoli Liu, Rui Zhang, Chang Liu et al.CVPR 2020
- Data-Free Network Compression via Parametric Non-uniform Mixed Precision QuantizationVladimir Chikin, Mikhail AntiukhCVPR 2022 · 17 citations
