AdaptiveNet: Post-deployment Neural Architecture Adaptation for Diverse Edge Environments
Hao Wen, Yuanchun Li, Zunshuai Zhang, Shiqi Jiang, Xiaozhou Ye, Ye Ouyang, Yaqin Zhang, Yunxin Liu
Abstract
Deep learning models are increasingly deployed to edge devices for real-time applications. To ensure stable service quality across diverse edge environments, it is highly desirable to generate tailored model architectures for different conditions. However, conventional pre-deployment model generation approaches are not satisfactory due to the difficulty of handling the diversity of edge environments and the demand for edge information. In this paper, we propose to adapt the model architecture after deployment in the target environment, where the model quality can be precisely measured and private edge data can be retained. To achieve efficient and effective edge model generation, we introduce a pretraining-assisted on-cloud model elastification method and an edge-friendly on-device architecture search method. Model elastification generates a high-quality search space of model architectures with the guidance of a developer-specified oracle model. Each subnet in the space is a valid model with different environment affinity, and each device efficiently finds and maintains the most suitable subnet based on a series of edge-tailored optimizations. Extensive experiments on various edge devices demonstrate that our approach is able to achieve significantly better accuracy-latency tradeoffs (e.g. 46.74% higher on average accuracy with a 60% latency budget) than strong baselines with minimal overhead (13 GPU hours in the cloud and 2 minutes on the edge server).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge DevicesXiangyu Li, Yuanchun Li, Yuanzhe Li, Ting Cao et al.MobiCom 2024 · 40 citations
- Multimodal Federated Learning via Contrastive Representation EnsembleQiying Yu, Yang Liu, Yimu Wang, Ke Xu et al.ICLR 2023 · 35 citations
- SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory BudgetRui Kong, Yuanchun Li, Qingtian Feng, Weijun Wang et al.ACL 2024 · 12 citations
- ContrastSense: Domain-invariant Contrastive Learning for In-the-Wild Wearable SensingGaole Dai, Huatao Xu, Hyungjun Yoon, Mo Li et al.UbiComp 2025 · 11 citations
- Elastic On-Device LLM ServiceWangsong Yin, Rongjie Yi, Daliang Xu, Gang Huang et al.MobiCom 2025 · 5 citations
Builds on20
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Data-Free Knowledge Distillation for Heterogeneous Federated LearningZhuangdi Zhu, Junyuan Hong, Jiayu ZhouICML 2021 · 957 citations
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 444 citations
Related papers
- End-to-End Model Generation with Large Language Models for Adaptive IoT Application DeploymentZhenyu Wen, Jintao Feng, Nanjie Yao, Di Wu et al.ICSE 2026
- Towards Robust and Efficient Cloud-Edge Elastic Model Adaptation via Selective Entropy DistillationYaofo Chen, Shuaicheng Niu, Yaowei Wang, Shoukai Xu et al.ICLR 2024 · 18 citations
- Mistify: Automating DNN Model Porting for On-Device Inference at the EdgePeizhen Guo, Bo Hu, Wenjun HuNSDI 2021 · 69 citations
- ZeroBN: Learning Compact Neural Networks For Latency-Critical Edge SystemsShuo Huai, Lei Zhang, Di Liu, Weichen Liu et al.DAC 2021 · 16 citations
- LitePred: Transferable and Scalable Latency Prediction for Hardware-Aware Neural Architecture SearchChengquan Feng, Li Lyna Zhang, Yuanchi Liu, Jiahang Xu et al.NSDI 2024 · 7 citations
