Mandheling: mixed-precision on-device DNN training with DSP offloading
Daliang Xu, Mengwei Xu, Qipeng Wang, Shangguang Wang, Yun Ma, Kang Huang, Gang Huang, Xin Jin, Xuanzhe Liu
Abstract
This paper proposes Mandheling, the first system that enables highly resource-efficient on-device training by orchestrating mixed-precision training with on-chip Digital Signal Processor (DSP) offloading. Mandheling fully explores the advantages of DSP in integer-based numerical calculations using four novel techniques: (1) a CPU-DSP co-scheduling scheme to situationally mitigate the overhead from DSP-unfriendly operators; (2) a self-adaptive rescaling algorithm to reduce the overhead of dynamic rescaling in backward propagation; (3) a batch-splitting algorithm to improve DSP cache efficiency; (4) a DSP compute subgraph-reusing mechanism to eliminate the preparation overhead on DSP. We have fully implemented Mandheling and demonstrated its effectiveness through extensive experiments. The results show that, compared to the state-of-the-art DNN engines from TFLite and MNN, Mandheling reduces per-batch training time by 5.5X and energy consumption by 8.9X on average. In end-to-end training tasks, Mandheling reduces convergence time by up to 10.7X and energy consumption by 13.1X, with only 1.9%--2.7% accuracy loss compared to the FP32 precision setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d40c86f5-c525-4d19-84c5-23e15cb02448Cited by top-tier papers16
- FwdLLM: Efficient Federated Finetuning of Large Language Models with Perturbed InferencesMengwei Xu, Dongqi Cai, Yaozong Wu, Xiang Li et al.USENIX ATC 2024 · 78 citations
- Fast On-device LLM Inference with NPUsDaliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu et al.ASPLOS 2025 · 38 citations
- Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge DevicesShengyuan Ye, Liekang Zeng, Xiaowen Chu, Guoliang Xing et al.MobiCom 2024 · 29 citations
- Federated Few-Shot Learning for Mobile NLPDongqi Cai, Shangguang Wang, Yaozong Wu, Felix Xiaozhu Lin et al.MobiCom 2023 · 28 citations
- TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce EdgeYoung D. Kwon, Rui Li, Stylianos I. Venieris, Jagmohan Chauhan et al.ICML 2024 · 25 citations
Builds on11
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Characterizing Impacts of Heterogeneity in Federated Learning upon Large-Scale Smartphone DataChengxu Yang, Qipeng Wang, Mengwei Xu, Zhenpeng Chen et al.WWW 2021 · 171 citations
- Hermes: an efficient federated learning framework for heterogeneous mobile clientsAng Li, Jingwei Sun, Pengcheng Li, Yu Pu et al.MobiCom 2021 · 167 citations
- NEMO: enabling neural-enhanced video streaming on commodity mobile devicesHyunho Yeo, Chan Ju Chong, Youngmok Jung, Juncheol Ye et al.MobiCom 2020 · 118 citations
- Billion-scale federated learning on mobile clients: a submodel design with tunable privacyChaoyue Niu, Fan Wu, Shaojie Tang, Lifeng Hua et al.MobiCom 2020 · 114 citations
Related papers
- Multi-Precision Policy Enforced Training (MuPPET) : A Precision-Switching Strategy for Quantised Fixed-Point Training of CNNsAditya Rajagopal, Diederik Adriaan Vink, Stylianos I. Venieris, Christos-Savvas BouganisICML 2020 · 17 citations
- Flex: Fast, Accurate DNN Inference on Low-Cost Edges Using Heterogeneous Accelerator ExecutionTanmoy Sen, Haiying Shen, Anand Padmanabha IyerEuroSys 2025 · 2 citations
- DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer ArithmeticHazem Hesham Yousef Shalby, Fabrizio Pittorino, Francesca Palermo, Diana Trojaniello et al.AAAI 2026 · 2 citations
- Campo: Cost-Aware Performance Optimization for Mixed-Precision Neural Network TrainingXin He, Jianhua Sun, Hao Chen, Dong LiUSENIX ATC 2022 · 10 citations
- Towards Energy-efficient Federated Learning via INT8-based Training on Mobile DSPsJinliang Yuan, Shangguang Wang, Hongyu Li, Daliang Xu et al.WWW 2024 · 8 citations
