RTMobile: Beyond Real-Time Mobile Acceleration of RNNs for Speech Recognition
Peiyan Dong, Siyue Wang, Wei Niu, Chengming Zhang, Sheng Lin, Zhengang Li, Yifan Gong, Bin Ren, Xue Lin, Dingwen Tao
Abstract
Recurrent neural networks (RNNs) based automatic speech recognition has nowadays become promising and important on mobile devices such as smart phones. However, previous RNN compression techniques either suffer from hardware performance overhead due to irregularity or significant accuracy loss due to the preserved regularity for hardware friendliness. In this work, we propose RTMobile that leverages both a novel block-based pruning approach and compiler optimizations to accelerate RNN inference on mobile devices. Our proposed RTMobile is the first work that can achieve real-time RNN inference on mobile platforms. Experimental results demonstrate that RTMobile can significantly outperform existing RNN hardware acceleration methods in terms of both inference accuracy and time. Compared with prior work on FPGA, RTMobile using Adreno 640 embedded GPU on GRU can improve the energy-efficiency by 40× while maintaining the same inference time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4aabfe5-5871-4a11-8420-9b3988ef67f0Cited by top-tier papers9
- DNNFusion: accelerating deep neural networks execution with advanced operator fusionWei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal et al.PLDI 2021 · 166 citations
- MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the EdgeGeng Yuan, Xiaolong Ma, Wei Niu, Zhengang Li et al.NeurIPS 2021 · 124 citations
- SparCL: Sparse Continual Learning on the EdgeZifeng Wang, Zheng Zhan, Yifan Gong, Geng Yuan et al.NeurIPS 2022 · 97 citations
- Achieving on-Mobile Real-Time Super-Resolution with Neural Architecture and Pruning SearchZheng Zhan, Yifan Gong, Pu Zhao, Geng Yuan et al.ICCV 2021 · 60 citations
- HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural NetworksJinqi Xiao, Chengming Zhang, Yu Gong, Miao Yin et al.AAAI 2023 · 35 citations
Related papers
- RT3D: Achieving Real-Time Execution of 3D Convolutional Neural Networks on Mobile DevicesWei Niu, Mengshu Sun, Zhengang Li, Jou-An Chen et al.AAAI 2021 · 14 citations
- Romou: rapidly generate high-performance tensor kernels for mobile GPUsRendong Liang, Ting Cao, Jicheng Wen, Manni Wang et al.MobiCom 2022 · 13 citations
- Accelerating Linear Recurrent Neural Networks for the Edge with Unstructured SparsityAlessandro Pierro, Steven Abreu, Jonathan Timcheck, Philipp Stratmann et al.ICML 2025
- Scalable Multi-FPGA Acceleration for Large RNNs with Full Parallelism LevelsDongup Kwon, Suyeon Hur, Hamin Jang, Eriko Nurvitadhi et al.DAC 2020 · 6 citations
- Extending the RISC-V ISA for Efficient RNN-based 5G Radio Resource ManagementRenzo Andri, Tomas Henriksson, Luca BeniniDAC 2020 · 10 citations
