Lune

INFOCOM2025顶会

Latency-Constrained Input-Aware Quantization of Time Series Inference Workflows at the Edge

Ashitabh Misra, Nurani Saoda, Tarek F. Abdelzaher

2025年份
1被引次数
1顶会引用

摘要

Quantization is a common technique for compressing Deep Neural Networks (DNNs) for edge device deployment, crucial for real-time time-series sensor data classification. While Adaptive Quantization (AQ) techniques allow updating the model's quantization scheme based on resource constraints without re-training, they are input-agnostic. Conversely, Input-aware Quantization (IQ) techniques adjust quantization per input, improving accuracy but ignoring resource constraints. Additionally, most quantization methods are designed for vision-based applications and underutilize time-frequency domain semantics. To this end, we present ReactQuant, an end-to-end framework for input-aware quantization of DNNs with strict resource constraints, specialized for time-series applications. ReactQuant features Partition-Derived Quantization, which segments time-series input based on its time and frequency domain features; a two-staged training process that restricts quantization space to each layer's promising bit-width candidates; and a novel algorithm to perform input-aware quantization within user-defined resource constraints. Extensive experimentation shows that ReactQuant outperforms AQ and IQ techniques in accuracy and ensures strict compliance with user-defined resource constraints. Under similar resource constraints, ReactQuant achieves up to 19% higher accuracy than AQ techniques and 4.6% higher than state-of-the-art IQ technique. It incurs no resource constraint violations, compared to up to 27.63% for state-of-the-art IQ technique.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖