StreamNet: Memory-Efficient Streaming Tiny Deep Learning Inference on the Microcontroller
Hong-Sheng Zheng, Yu-Yuan Liu, Chen-Fong Hsu, Tsung Tai Yeh
Abstract
With the emerging Tiny Machine Learning (TinyML) inference applications, there is a growing interest when deploying TinyML models on the low-power Microcon-troller Unit (MCU). However, deploying TinyML models on MCUs reveals several challenges due to the MCU’s resource constraints, such as small flash memory, tight SRAM memory budget, and slow CPU performance. Unlike typical layer-wise inference, patch-based inference reduces the peak usage of SRAM memory on MCUs by saving small patches rather than the entire tensor in the SRAM memory. However, the processing of patch-based inference tremendously increases the amount of MACs against the layer-wise method. Thus, this notoriously computational overhead makes patch-based inference undesirable on MCUs. This work designs StreamNet that employs the stream buffer to eliminate the redundant computation of patch-based inference. StreamNet uses 1D and 2D streaming processing and provides an parameter selection algorithm that automatically improve the performance of patch-based inference with minimal requirements on the MCU’s SRAM memory space. In 10 TinyML models, StreamNet-2D achieves a geometric mean of 7.3X speedup and saves 81% of MACs over the state-of-the-art patch-based inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bbc6a93-1df8-4eed-8df4-4854af12ef73Cited by top-tier papers2
- DEX: Data Channel Extension for Efficient CNN Inference on Tiny AI AcceleratorsTaesik Gong, Fahim Kawsar, Chulhong MinNeurIPS 2024 · 8 citations
- msf-CNN: Patch-based Multi-Stage Fusion with Convolutional Neural Networks for TinyMLZhaolan Huang, Emmanuel BaccelliNeurIPS 2025 · 4 citations
Builds on3
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn et al.NeurIPS 2020 · 827 citations
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- RNNPool: Efficient Non-linear Pooling for RAM Constrained InferenceOindrila Saha, Aditya Kusupati, Harsha Vardhan Simhadri, Manik Varma et al.NeurIPS 2020 · 58 citations
Related papers
- TinyTS: Memory-Efficient TinyML Model Compiler Framework on MicrocontrollersYu-Yuan Liu, Hong-Sheng Zheng, Yu Fang Hu, Chen-Fong Hsu et al.HPCA 2024 · 10 citations
- Memory-efficient Patch-based Inference for Tiny Deep LearningJi Lin, Wei-Ming Chen, Han Cai, Chuang Gan et al.NeurIPS 2021 · 190 citations
- AtomNet: Designing Tiny Models from Operators Under Extreme MCU ConstraintsZhiwei Dong, Mingzhu Shen, Shihao Bai, Xiuying Wei et al.AAAI 2025
- IP Protection in TinyMLJinwen Wang, Yuhao Wu, Han Liu, Bo Yuan et al.DAC 2023 · 6 citations
- DTMM: Deploying TinyML Models on Extremely Weak IoT Devices with PruningLixiang Han, Zhen Xiao, Zhenjiang LiINFOCOM 2024 · 20 citations
