Dancing along Battery: Enabling Transformer with Run-time Reconfigurability on Mobile Devices
Yuhong Song, Weiwen Jiang, Bingbing Li, Panjie Qi, Qingfeng Zhuge, Edwin Hsing-Mean Sha, Sakyasingha Dasgupta, Yiyu Shi, Caiwen Ding
摘要
A pruning-based AutoML framework for run-time reconfigurability, namely RT 3 , is proposed in this work. This enables Transformer-based large Natural Language Processing (NLP) models to be efficiently executed on resource-constrained mobile devices and reconfigured (i.e., switching models for dynamic hardware conditions) at run-time. Such reconfigurability is the key to save energy for battery-powered mobile devices, which widely use dynamic voltage and frequency scaling (DVFS) technique for hardware reconfiguration to prolong battery life. In this work, we creatively explore a hybrid block-structured pruning (BP) and pattern pruning (PP) for Transformer-based models and first attempt to combine hardware and software reconfiguration to maximally save energy for battery-powered mobile devices. Specifically, RT 3 integrates two-level optimizations: First, it utilizes an efficient BP as the first-step compression for resourceconstrained mobile devices; then, RT 3 heuristically generates a shrunken search space based on the first level optimization and searches multiple pattern sets with diverse sparsity for PP via reinforcement learning to support lightweight software reconfiguration, which corresponds to available frequency levels of DVFS (i.e., hardware reconfiguration). At run-time, RT 3 can switch the lightweight pattern sets within 45ms to guarantee the required real-time constraint at different frequency levels. Results further show that RT 3 can prolong battery life over 4× improvement with less than 1% accuracy loss for Transformer and 1.5% score decrease for DistilBERT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai 等ACL 2020 · 被引用 215 次
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang 等ASPLOS 2020 · 被引用 214 次
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesXiaolong Ma, Fu-Ming Guo, Wei Niu, Xue Lin 等AAAI 2020 · 被引用 201 次
- Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple TasksLei Yang, Zheyu Yan, Meng Li, Hyoukjun Kwon 等DAC 2020 · 被引用 115 次
- Non-uniform DNN Structured Subnets Sampling for Dynamic InferenceLi Yang, Zhezhi He, Yu Cao, Deliang FanDAC 2020 · 被引用 12 次
相关 Paper
- RL-PTQ: RL-based Mixed Precision Quantization for Hybrid Vision TransformersEunji Kwon, Minxuan Zhou, Weihong Xu, Tajana Rosing 等DAC 2024 · 被引用 4 次
- RTMobile: Beyond Real-Time Mobile Acceleration of RNNs for Speech RecognitionPeiyan Dong, Siyue Wang, Wei Niu, Chengming Zhang 等DAC 2020 · 被引用 50 次
- ASBP: Automatic Structured Bit-Pruning for RRAM-based NN AcceleratorSongyun Qu, Bing Li, Ying Wang, Lei ZhangDAC 2021 · 被引用 15 次
- NPAS: A Compiler-Aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile AccelerationZhengang Li, Geng Yuan, Wei Niu, Pu Zhao 等CVPR 2021
- Storage Efficient and Dynamic Flexible Runtime Channel Pruning via Deep Reinforcement LearningJianda Chen, Shangyu Chen, Sinno Jialin PanNeurIPS 2020 · 被引用 31 次
