Mandheling: mixed-precision on-device DNN training with DSP offloading
Daliang Xu, Mengwei Xu, Qipeng Wang, Shangguang Wang, Yun Ma, Kang Huang, Gang Huang, Xin Jin, Xuanzhe Liu
摘要
This paper proposes Mandheling, the first system that enables highly resource-efficient on-device training by orchestrating mixed-precision training with on-chip Digital Signal Processor (DSP) offloading. Mandheling fully explores the advantages of DSP in integer-based numerical calculations using four novel techniques: (1) a CPU-DSP co-scheduling scheme to situationally mitigate the overhead from DSP-unfriendly operators; (2) a self-adaptive rescaling algorithm to reduce the overhead of dynamic rescaling in backward propagation; (3) a batch-splitting algorithm to improve DSP cache efficiency; (4) a DSP compute subgraph-reusing mechanism to eliminate the preparation overhead on DSP. We have fully implemented Mandheling and demonstrated its effectiveness through extensive experiments. The results show that, compared to the state-of-the-art DNN engines from TFLite and MNN, Mandheling reduces per-batch training time by 5.5X and energy consumption by 8.9X on average. In end-to-end training tasks, Mandheling reduces convergence time by up to 10.7X and energy consumption by 13.1X, with only 1.9%--2.7% accuracy loss compared to the FP32 precision setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- FwdLLM: Efficient Federated Finetuning of Large Language Models with Perturbed InferencesMengwei Xu, Dongqi Cai, Yaozong Wu, Xiang Li 等USENIX ATC 2024 · 被引用 78 次
- Fast On-device LLM Inference with NPUsDaliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu 等ASPLOS 2025 · 被引用 38 次
- Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge DevicesShengyuan Ye, Liekang Zeng, Xiaowen Chu, Guoliang Xing 等MobiCom 2024 · 被引用 29 次
- Federated Few-Shot Learning for Mobile NLPDongqi Cai, Shangguang Wang, Yaozong Wu, Felix Xiaozhu Lin 等MobiCom 2023 · 被引用 28 次
- TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce EdgeYoung D. Kwon, Rui Li, Stylianos I. Venieris, Jagmohan Chauhan 等ICML 2024 · 被引用 25 次
它引用的顶会 Paper11
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett 等ICLR 2021 · 被引用 1,917 次
- Characterizing Impacts of Heterogeneity in Federated Learning upon Large-Scale Smartphone DataChengxu Yang, Qipeng Wang, Mengwei Xu, Zhenpeng Chen 等WWW 2021 · 被引用 171 次
- Hermes: an efficient federated learning framework for heterogeneous mobile clientsAng Li, Jingwei Sun, Pengcheng Li, Yu Pu 等MobiCom 2021 · 被引用 167 次
- NEMO: enabling neural-enhanced video streaming on commodity mobile devicesHyunho Yeo, Chan Ju Chong, Youngmok Jung, Juncheol Ye 等MobiCom 2020 · 被引用 118 次
- Billion-scale federated learning on mobile clients: a submodel design with tunable privacyChaoyue Niu, Fan Wu, Shaojie Tang, Lifeng Hua 等MobiCom 2020 · 被引用 114 次
相关 Paper
- Multi-Precision Policy Enforced Training (MuPPET) : A Precision-Switching Strategy for Quantised Fixed-Point Training of CNNsAditya Rajagopal, Diederik Adriaan Vink, Stylianos I. Venieris, Christos-Savvas BouganisICML 2020 · 被引用 17 次
- Flex: Fast, Accurate DNN Inference on Low-Cost Edges Using Heterogeneous Accelerator ExecutionTanmoy Sen, Haiying Shen, Anand Padmanabha IyerEuroSys 2025 · 被引用 2 次
- DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer ArithmeticHazem Hesham Yousef Shalby, Fabrizio Pittorino, Francesca Palermo, Diana Trojaniello 等AAAI 2026 · 被引用 2 次
- Campo: Cost-Aware Performance Optimization for Mixed-Precision Neural Network TrainingXin He, Jianhua Sun, Hao Chen, Dong LiUSENIX ATC 2022 · 被引用 10 次
- Towards Energy-efficient Federated Learning via INT8-based Training on Mobile DSPsJinliang Yuan, Shangguang Wang, Hongyu Li, Daliang Xu 等WWW 2024 · 被引用 8 次
