Dancing along Battery: Enabling Transformer with Run-time Reconfigurability on Mobile Devices
Yuhong Song, Weiwen Jiang, Bingbing Li, Panjie Qi, Qingfeng Zhuge, Edwin Hsing-Mean Sha, Sakyasingha Dasgupta, Yiyu Shi, Caiwen Ding
Abstract
A pruning-based AutoML framework for run-time reconfigurability, namely RT 3 , is proposed in this work. This enables Transformer-based large Natural Language Processing (NLP) models to be efficiently executed on resource-constrained mobile devices and reconfigured (i.e., switching models for dynamic hardware conditions) at run-time. Such reconfigurability is the key to save energy for battery-powered mobile devices, which widely use dynamic voltage and frequency scaling (DVFS) technique for hardware reconfiguration to prolong battery life. In this work, we creatively explore a hybrid block-structured pruning (BP) and pattern pruning (PP) for Transformer-based models and first attempt to combine hardware and software reconfiguration to maximally save energy for battery-powered mobile devices. Specifically, RT 3 integrates two-level optimizations: First, it utilizes an efficient BP as the first-step compression for resourceconstrained mobile devices; then, RT 3 heuristically generates a shrunken search space based on the first level optimization and searches multiple pattern sets with diverse sparsity for PP via reinforcement learning to support lightweight software reconfiguration, which corresponds to available frequency levels of DVFS (i.e., hardware reconfiguration). At run-time, RT 3 can switch the lightweight pattern sets within 45ms to guarantee the required real-time constraint at different frequency levels. Results further show that RT 3 can prolong battery life over 4× improvement with less than 1% accuracy loss for Transformer and 1.5% score decrease for DistilBERT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on5
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai et al.ACL 2020 · 215 citations
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang et al.ASPLOS 2020 · 214 citations
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesXiaolong Ma, Fu-Ming Guo, Wei Niu, Xue Lin et al.AAAI 2020 · 201 citations
- Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple TasksLei Yang, Zheyu Yan, Meng Li, Hyoukjun Kwon et al.DAC 2020 · 115 citations
- Non-uniform DNN Structured Subnets Sampling for Dynamic InferenceLi Yang, Zhezhi He, Yu Cao, Deliang FanDAC 2020 · 12 citations
Related papers
- RL-PTQ: RL-based Mixed Precision Quantization for Hybrid Vision TransformersEunji Kwon, Minxuan Zhou, Weihong Xu, Tajana Rosing et al.DAC 2024 · 4 citations
- RTMobile: Beyond Real-Time Mobile Acceleration of RNNs for Speech RecognitionPeiyan Dong, Siyue Wang, Wei Niu, Chengming Zhang et al.DAC 2020 · 50 citations
- ASBP: Automatic Structured Bit-Pruning for RRAM-based NN AcceleratorSongyun Qu, Bing Li, Ying Wang, Lei ZhangDAC 2021 · 15 citations
- NPAS: A Compiler-Aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile AccelerationZhengang Li, Geng Yuan, Wei Niu, Pu Zhao et al.CVPR 2021
- Storage Efficient and Dynamic Flexible Runtime Channel Pruning via Deep Reinforcement LearningJianda Chen, Shangyu Chen, Sinno Jialin PanNeurIPS 2020 · 31 citations
