AdaptQNet: Optimizing Quantized DNN on Microcontrollers via Adaptive Heterogeneous Processing Unit Utilization
Yansong Sun, Jialuo He, Dirk Kutscher, Huangxun Chen
Abstract
There is a growing trend in deploying DNNs on tiny microcontroller (MCUs) to provide inference capabilities in the IoT. While prior research has explored many lightweight techniques to compress DNN models, achieving overall efficiency in model inference requires not only model optimization but also careful system resource utilization for execution. Existing studies primarily leverage arithmetic logic units (ALUs) for integer-only computations on a single CPU core. Floating-point units (FPU) and multi-core capabilities available in many existing MCUs remain underutilized. To fill this gap, we propose AdaptQNet, a novel MCU neural network system that can determine the optimal precision assignment for different layers of a DNN model. AdaptQNet models the latency of various operators in DNN models across different precisions on heterogeneous processing units. This facilitates the discovery of models that utilize FPU and multi-core capabilities to enhance capacity while adhering to stringent memory constraints. Our implementation and experiments demonstrate that AdaptQNet enables the deployment of models with better accuracy-efficiency trade-off on MCUs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 87e5660f-43aa-476e-98cd-77f246ddec77Related papers
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn et al.NeurIPS 2020 · 827 citations
- InstantNet: Automated Generation and Deployment of Instantaneously Switchable-Precision NetworksYonggan Fu, Zhongzhi Yu, Yongan Zhang, Yifan Jiang et al.DAC 2021 · 6 citations
- UDC: Unified DNAS for Compressible TinyML Models for Neural Processing UnitsIgor Fedorov, Ramon Matas Navarro, Hokchhay Tann, Chuteng Zhou et al.NeurIPS 2022 · 19 citations
- Intermittent Inference with Nonuniformly Compressed Multi-Exit Neural Network for Energy Harvesting Powered DevicesYawen Wu, Zhepeng Wang, Zhenge Jia, Yiyu Shi et al.DAC 2020 · 41 citations
- Flex: Fast, Accurate DNN Inference on Low-Cost Edges Using Heterogeneous Accelerator ExecutionTanmoy Sen, Haiying Shen, Anand Padmanabha IyerEuroSys 2025 · 2 citations
