Context-aware Adaptive Surgery: A Fast and Effective Framework for Adaptative Model Partition
Hongli Wang, Bin Guo, Jiaqi Liu, Sicong Liu, Yungang Wu, Zhiwen Yu
Abstract
Deep Neural Networks (DNNs) have made massive progress in many fields and deploying DNNs on end devices has become an emerging trend to make intelligence closer to users. However, it is challenging to deploy large-scale and computation-intensive DNNs on resource-constrained end devices due to their small size and lightweight. To this end, model partition, which aims to partition DNNs into multiple parts to realize the collaborative computing of multiple devices, has received extensive research attention. To find the optimal partition, most existing approaches need to run from scratch under given resource constraints. However, they ignore that resources of devices (e.g., storage, battery power), and performance requirements (e.g., inference latency), are often continuously changing, making the optimal partition solution change constantly during processing. Therefore, it is very important to reduce the tuning latency of model partition to realize the real-time adaption under the changing processing context. To address these problems, we propose the Context-aware Adaptive Surgery (CAS) framework to actively perceive the changing processing context, and adaptively find the appropriate partition solution in real-time. Specifically, we construct the partition state graph to comprehensively model different partition solutions of DNNs by import context resources. Then "the neighbor effect" is proposed, which provides the heuristic rule for the search process. When the processing context changes, CAS adopts the runtime search algorithm, Graph-based Adaptive DNN Surgery (GADS), to quickly find the appropriate partition that satisfies resource constraints under the guidance of the neighbor effect. The experimental results show that CAS realizes adaptively rapid tuning of the model partition solutions in 10ms scale even for large DNNs (2.25x to 221.7x search time improvement than the state-of-the-art researches), and the total inference latency still keeps the same level with baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2df71f85-f51a-4739-a50e-61caa4dcc68fRelated papers
- Autodidactic Neurosurgeon: Collaborative Deep Inference for Mobile Edge Intelligence via Online LearningLetian Zhang, Lixing Chen, Jie XuWWW 2021 · 75 citations
- Towards Real-time Cooperative Deep Inference over the Cloud and Edge End DevicesShigeng Zhang, Yinggang Li, Xuan Liu, Song Guo et al.UbiComp 2020 · 77 citations
- End-to-End Model Generation with Large Language Models for Adaptive IoT Application DeploymentZhenyu Wen, Jintao Feng, Nanjie Yao, Di Wu et al.ICSE 2026
- MoteNN: Memory Optimization via Fine-grained Scheduling for Deep Neural Networks on Tiny DevicesRenze Chen, Zijian Ding, Size Zheng, Meng Li et al.DAC 2024 · 7 citations
- DeepAdapter: A Collaborative Deep Learning Framework for the Mobile Web Using Context-Aware Network PruningYakun Huang, Xiuquan Qiao, Jian Tang, Pei Ren et al.INFOCOM 2020 · 32 citations
