Shiftry: RNN inference in 2KB of RAM
Aayan Kumar, Vivek Seshadri, Rahul Sharma
Abstract
Traditionally, IoT devices send collected sensor data to an intelligent cloud where machine learning (ML) inference happens. However, this course is rapidly changing and there is a recent trend to run ML on the edge IoT devices themselves. An intelligent edge is attractive because it saves network round trip (efficiency) and keeps user data at the source (privacy). However, the IoT devices are much more resource constrained than the cloud, which makes running ML on them challenging. Specifically, consider Arduino Uno, a commonly used board, that has 2KB of RAM and 32KB of read-only Flash memory. Although recent breakthroughs in ML have created novel recurrent neural network (RNN) models that provide good accuracy with KB-sized models, deploying them on tiny devices with such hard memory requirements has remained elusive.
We provide, Shiftry, an automatic compiler from high-level floating-point ML models to fixed-point C-programs with 8-bit and 16-bit integers, which have significantly lower memory requirements. For this conversion, Shiftry uses a data-driven float-to-fixed procedure and a RAM management mechanism. These techniques enable us to provide first empirical evaluation of RNNs running on tiny edge devices. On simpler ML models that prior work could handle, Shiftry-generated code has lower latency and higher accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 062541e3-308d-42d7-8913-023ff094c723Cited by top-tier papers3
- SiRnn: A Math Library for Secure RNN InferenceDeevashwer Rathee, Mayank Rathee, Rahul Kranti Kiran Goli, Divya Gupta et al.S&P 2021 · 154 citations
- LiteFlow: towards high-performance adaptive neural networks for kernel datapathJunxue Zhang, Chaoliang Zeng, Hong Zhang, Shuihai Hu et al.SIGCOMM 2022 · 23 citations
- Cost of Soundness in Mixed-Precision TuningAnastasia Isychev, Debasmita LoharOOPSLA 2025 · 2 citations
Builds on2
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 622 citations
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
Related papers
- HiRISE: High-Resolution Image Scaling for Edge ML via In-Sensor Compression and Selective ROIBrendan Reidy, Sepehr Tabrizchi, Mohammadreza Mohammadi, Shaahin Angizi et al.DAC 2024 · 3 citations
- Input-Dependent Edge-Cloud Mapping of Recurrent Neural Networks InferenceDaniele Jahier Pagliari, Roberta Chiaro, Yukai Chen, Sara Vinco et al.DAC 2020 · 9 citations
- TinyTTA: Efficient Test-time Adaptation via Early-exit Ensembles on Edge DevicesHong Jia, Young D. Kwon, Alessio Orsino, Ting Dang et al.NeurIPS 2024 · 25 citations
- MicroVSA: An Ultra-Lightweight Vector Symbolic Architecture-based Classifier Library for Always-On Inference on Tiny MicrocontrollersNuntipat Narkthong, Shijin Duan, Shaolei Ren, Xiaolin XuASPLOS 2024 · 8 citations
- Entropy-Driven Mixed-Precision Quantization for Deep Network DesignZhenhong Sun, Ce Ge, Junyan Wang, Ming Lin et al.NeurIPS 2022 · 41 citations
