RTMobile: Beyond Real-Time Mobile Acceleration of RNNs for Speech Recognition
Peiyan Dong, Siyue Wang, Wei Niu, Chengming Zhang, Sheng Lin, Zhengang Li, Yifan Gong, Bin Ren, Xue Lin, Dingwen Tao
摘要
Recurrent neural networks (RNNs) based automatic speech recognition has nowadays become promising and important on mobile devices such as smart phones. However, previous RNN compression techniques either suffer from hardware performance overhead due to irregularity or significant accuracy loss due to the preserved regularity for hardware friendliness. In this work, we propose RTMobile that leverages both a novel block-based pruning approach and compiler optimizations to accelerate RNN inference on mobile devices. Our proposed RTMobile is the first work that can achieve real-time RNN inference on mobile platforms. Experimental results demonstrate that RTMobile can significantly outperform existing RNN hardware acceleration methods in terms of both inference accuracy and time. Compared with prior work on FPGA, RTMobile using Adreno 640 embedded GPU on GRU can improve the energy-efficiency by 40× while maintaining the same inference time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- DNNFusion: accelerating deep neural networks execution with advanced operator fusionWei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal 等PLDI 2021 · 被引用 166 次
- MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the EdgeGeng Yuan, Xiaolong Ma, Wei Niu, Zhengang Li 等NeurIPS 2021 · 被引用 124 次
- SparCL: Sparse Continual Learning on the EdgeZifeng Wang, Zheng Zhan, Yifan Gong, Geng Yuan 等NeurIPS 2022 · 被引用 97 次
- Achieving on-Mobile Real-Time Super-Resolution with Neural Architecture and Pruning SearchZheng Zhan, Yifan Gong, Pu Zhao, Geng Yuan 等ICCV 2021 · 被引用 60 次
- HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural NetworksJinqi Xiao, Chengming Zhang, Yu Gong, Miao Yin 等AAAI 2023 · 被引用 35 次
相关 Paper
- RT3D: Achieving Real-Time Execution of 3D Convolutional Neural Networks on Mobile DevicesWei Niu, Mengshu Sun, Zhengang Li, Jou-An Chen 等AAAI 2021 · 被引用 14 次
- Romou: rapidly generate high-performance tensor kernels for mobile GPUsRendong Liang, Ting Cao, Jicheng Wen, Manni Wang 等MobiCom 2022 · 被引用 13 次
- Accelerating Linear Recurrent Neural Networks for the Edge with Unstructured SparsityAlessandro Pierro, Steven Abreu, Jonathan Timcheck, Philipp Stratmann 等ICML 2025
- Scalable Multi-FPGA Acceleration for Large RNNs with Full Parallelism LevelsDongup Kwon, Suyeon Hur, Hamin Jang, Eriko Nurvitadhi 等DAC 2020 · 被引用 6 次
- Extending the RISC-V ISA for Efficient RNN-based 5G Radio Resource ManagementRenzo Andri, Tomas Henriksson, Luca BeniniDAC 2020 · 被引用 10 次
