FracTrain: Fractionally Squeezing Bit Savings Both Temporally and Spatially for Efficient DNN Training
Yonggan Fu, Haoran You, Yang Zhao, Yue Wang, Chaojian Li, Kailash Gopalakrishnan, Zhangyang Wang, Yingyan Lin
摘要
Recent breakthroughs in deep neural networks (DNNs) have fueled a tremendous demand for intelligent edge devices featuring on-site learning, while the practical realization of such systems remains a challenge due to the limited resources available at the edge and the required massive training costs for state-of-the-art (SOTA) DNNs. As reducing precision is one of the most effective knobs for boosting training time/energy efficiency, there has been a growing interest in low-precision DNN training. In this paper, we explore from an orthogonal direction: how to fractionally squeeze out more training cost savings from the most redundant bit level, progressively along the training trajectory and dynamically per input. Specifically, we propose FracTrain that integrates (i) progressive fractional quantization which gradually increases the precision of activations, weights, and gradients that will not reach the precision of SOTA static quantized DNN training until the final training stage, and (ii) dynamic fractional quantization which assigns precisions to both the activations and gradients of each layer in an input-adaptive manner, for only "fractionally" updating layer parameters. Extensive simulations and ablation studies (six models, four datasets, and three training settings including standard, adaptation, and fine-tuning) validate the effectiveness of FracTrain in reducing computational cost and hardware-quantified energy/latency of DNN training while achieving a comparable or better (-0.12% ∼ +1.87%) accuracy. For example, when training ResNet-74 on CIFAR-10, FracTrain achieves 77.6% and 53.5% computational cost and training latency savings, respectively, compared with the best SOTA baseline, while achieving a comparable (-0.07%) accuracy. Our codes are available at: https://github.com/RICE-EIC/FracTrain .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- DAdaQuant: Doubly-adaptive quantization for communication-efficient Federated LearningRobert Hönig, Yiren Zhao, Robert MullinsICML 2022 · 被引用 87 次
- Double-Win Quant: Aggressively Winning Robustness of Quantized Deep Neural Networks via Random Precision Training and InferenceYonggan Fu, Qixuan Yu, Meng Li, Vikas Chandra 等ICML 2021 · 被引用 34 次
- Auto-scaling Vision Transformers without TrainingWuyang Chen, Wei Huang, Xianzhi Du, Xiaodan Song 等ICLR 2022 · 被引用 27 次
- DepthShrinker: A New Compression Paradigm Towards Boosting Real-Hardware Efficiency of Compact Neural NetworksYonggan Fu, Haichuan Yang, Jiayi Yuan, Meng Li 等ICML 2022 · 被引用 26 次
- MIA-Former: Efficient and Robust Vision Transformers via Multi-Grained Input-AdaptationZhongzhi Yu, Yonggan Fu, Sicheng Li, Chaojian Li 等AAAI 2022 · 被引用 20 次
它引用的顶会 Paper5
- Drawing Early-Bird Tickets: Toward More Efficient Training of Deep NetworksHaoran You, Chaojian Li, Pengfei Xu, Yonggan Fu 等ICLR 2020 · 被引用 282 次
- DRQ: Dynamic Region-based Quantization for Deep Neural Network AccelerationZhuoran Song, Bangqi Fu, Feiyang Wu, Zhaoming Jiang 等ISCA 2020 · 被引用 92 次
- EDD: Efficient Differentiable DNN Architecture and Implementation Co-search for Embedded AI SolutionsYuhong Li, Cong Hao, Xiaofan Zhang, Xinheng Liu 等DAC 2020 · 被引用 79 次
- Fractional Skipping: Towards Finer-Grained Dynamic CNN InferenceJianghao Shen, Yue Wang, Pengfei Xu, Yonggan Fu 等AAAI 2020 · 被引用 49 次
- Towards Unified INT8 Training for Convolutional Neural NetworkFeng Zhu, Ruihao Gong, Fengwei Yu, Xianglong Liu 等CVPR 2020
相关 Paper
- Not All Bits have Equal Value: Heterogeneous Precisions via Trainable NoisePedro Savarese, Xin Yuan, Yanjing Li, Michael MaireNeurIPS 2022 · 被引用 9 次
- CPT: Efficient Deep Neural Network Training via Cyclic PrecisionYonggan Fu, Han Guo, Meng Li, Xin Yang 等ICLR 2021 · 被引用 36 次
- DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer ArithmeticHazem Hesham Yousef Shalby, Fabrizio Pittorino, Francesca Palermo, Diana Trojaniello 等AAAI 2026 · 被引用 2 次
- Towards Cheaper Inference in Deep Networks with Lower Bit-Width AccumulatorsYaniv Blumenfeld, Itay Hubara, Daniel SoudryICLR 2024 · 被引用 5 次
- Flex: Fast, Accurate DNN Inference on Low-Cost Edges Using Heterogeneous Accelerator ExecutionTanmoy Sen, Haiying Shen, Anand Padmanabha IyerEuroSys 2025 · 被引用 2 次
