Shiftry: RNN inference in 2KB of RAM
Aayan Kumar, Vivek Seshadri, Rahul Sharma
摘要
Traditionally, IoT devices send collected sensor data to an intelligent cloud where machine learning (ML) inference happens. However, this course is rapidly changing and there is a recent trend to run ML on the edge IoT devices themselves. An intelligent edge is attractive because it saves network round trip (efficiency) and keeps user data at the source (privacy). However, the IoT devices are much more resource constrained than the cloud, which makes running ML on them challenging. Specifically, consider Arduino Uno, a commonly used board, that has 2KB of RAM and 32KB of read-only Flash memory. Although recent breakthroughs in ML have created novel recurrent neural network (RNN) models that provide good accuracy with KB-sized models, deploying them on tiny devices with such hard memory requirements has remained elusive.
We provide, Shiftry, an automatic compiler from high-level floating-point ML models to fixed-point C-programs with 8-bit and 16-bit integers, which have significantly lower memory requirements. For this conversion, Shiftry uses a data-driven float-to-fixed procedure and a RAM management mechanism. These techniques enable us to provide first empirical evaluation of RNNs running on tiny edge devices. On simpler ML models that prior work could handle, Shiftry-generated code has lower latency and higher accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SiRnn: A Math Library for Secure RNN InferenceDeevashwer Rathee, Mayank Rathee, Rahul Kranti Kiran Goli, Divya Gupta 等S&P 2021 · 被引用 154 次
- LiteFlow: towards high-performance adaptive neural networks for kernel datapathJunxue Zhang, Chaoliang Zeng, Hong Zhang, Shuihai Hu 等SIGCOMM 2022 · 被引用 23 次
- Cost of Soundness in Mixed-Precision TuningAnastasia Isychev, Debasmita LoharOOPSLA 2025 · 被引用 2 次
它引用的顶会 Paper2
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 被引用 622 次
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li 等ICCV 2019 · 被引用 540 次
相关 Paper
- HiRISE: High-Resolution Image Scaling for Edge ML via In-Sensor Compression and Selective ROIBrendan Reidy, Sepehr Tabrizchi, Mohammadreza Mohammadi, Shaahin Angizi 等DAC 2024 · 被引用 3 次
- Input-Dependent Edge-Cloud Mapping of Recurrent Neural Networks InferenceDaniele Jahier Pagliari, Roberta Chiaro, Yukai Chen, Sara Vinco 等DAC 2020 · 被引用 9 次
- TinyTTA: Efficient Test-time Adaptation via Early-exit Ensembles on Edge DevicesHong Jia, Young D. Kwon, Alessio Orsino, Ting Dang 等NeurIPS 2024 · 被引用 25 次
- MicroVSA: An Ultra-Lightweight Vector Symbolic Architecture-based Classifier Library for Always-On Inference on Tiny MicrocontrollersNuntipat Narkthong, Shijin Duan, Shaolei Ren, Xiaolin XuASPLOS 2024 · 被引用 8 次
- Entropy-Driven Mixed-Precision Quantization for Deep Network DesignZhenhong Sun, Ce Ge, Junyan Wang, Ming Lin 等NeurIPS 2022 · 被引用 41 次
