Unlocking the Non-deterministic Computing Power with Memory-Elastic Multi-Exit Neural Networks
Jiaming Huang, Yi Gao, Wei Dong
Abstract
With the increasing demand for Web of Things (WoT) and edge computing, the efficient utilization of limited computing power on edge devices is becoming a crucial challenge. Traditional neural networks (NNs) as web services rely on deterministic computational resources. However, they may fail to output the results on non-deterministic computing power which could be preempted at any time, degrading the task performance significantly. Multi-exit NNs with multiple branches have been proposed as a solution, but the accuracy of intermediate results may be unsatisfactory. In this paper, we propose MEEdge, a system that automatically transforms classic single-exit models into heterogeneous and dynamic multi-exit models which enables Memory-Elastic inference at the Edge with non-deterministic computing power. To build heterogeneous multi-exit models, MEEdge uses efficient convolutions to form a branch zoo and High Priority First (HPF)-based branch placement method for branch growth. To adapt models to dynamically varying computational resources, we employ a novel on-device scheduler for collaboration. Further, to reduce the memory overhead caused by dynamic branches, we propose neuron-level weight sharing and few-shot knowledge distillation(KD) retraining. Our experimental results show that models generated by MEEdge can achieve up to 27.31% better performance than existing multi-exit NNs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Resource-aware Deployment of Dynamic DNNs over Multi-tiered Interconnected SystemsChetna Singhal, Yashuo Wu, Francesco Malandrino, Marco Levorato et al.INFOCOM 2024 · 15 citations
- Resilient and Communication Efficient Learning for Heterogeneous Federated SystemsZhuangdi Zhu, Junyuan Hong, Steve Drew, Jiayu ZhouICML 2022 · 46 citations
- FlexiFed: Personalized Federated Learning for Edge Clients with Heterogeneous Model ArchitecturesKaibin Wang, Qiang He, Feifei Chen, Chunyang Chen et al.WWW 2023 · 69 citations
- SIEVE: Speculative Inference on the Edge with Versatile ExportationBabak Zamirai, Salar Latifi, Pedram Zamirai, Scott A. MahlkeDAC 2020 · 5 citations
- Condense: A Framework for Device and Frequency Adaptive Neural Network Models on the EdgeYifan Gong, Pu Zhao, Zheng Zhan, Yushu Wu et al.DAC 2023 · 4 citations
