Condense: A Framework for Device and Frequency Adaptive Neural Network Models on the Edge
Yifan Gong, Pu Zhao, Zheng Zhan, Yushu Wu, Chao Wu, Zhenglun Kong, Minghai Qin, Caiwen Ding, Yanzhi Wang
摘要
With the popularity of battery-powered edge computing, an important yet under-explored problem is the supporting of DNNs for diverse edge devices. On the one hand, different edge platforms have various runtime requirements and computation/memory capabilities. Deploying the same DNN model is unsatisfiable, while designing a specialized DNN for each platform is prohibitively expensive. On the other hand, for a single edge device, DVFS is leveraged to prolong the battery, incurring significant inference speed variation for the same DNN and consequently poor user experience. To tackle this, we propose Condense, a framework providing a single adaptive model that can be reconfigured (switch to various sub-networks with different computations/parameters) instantly for diverse devices and execution frequencies without any retraining. Experiments demonstrate that Condense can simultaneously provide vast high-accuracy sub-networks with different computations and parameters corresponding to various sparsity ratios to support diverse edge devices with different runtime requirements, and reduce the speed variation under varying frequencies on each device, with a memory cost of only one set of weights.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Exploring Token Pruning in Vision State Space ModelsZheng Zhan, Zhenglun Kong, Yifan Gong, Yushu Wu 等NeurIPS 2024 · 被引用 34 次
- LOTUS: learning-based online thermal and latency variation management for two-stage detectors on edge devicesYifan Gong, Yushu Wu, Zheng Zhan, Pu Zhao 等DAC 2024 · 被引用 9 次
- Numerical Pruning for Efficient Autoregressive ModelsXuan Shen, Zhao Song, Yufa Zhou, Bo Chen 等AAAI 2025 · 被引用 5 次
- InfScaler: Enabling Efficient ML Inference Serving on Multi-Accelerator Edge Devices via Asymmetric Auto-ScalingBorui Li, Tiange Xia, Shuai Wang, Shuai WangDAC 2025 · 被引用 2 次
相关 Paper
- PowerPruning: Selecting Weights and Activations for Power-Efficient Neural Network AccelerationRichard Petri, Grace Li Zhang, Yiran Chen, Ulf Schlichtmann 等DAC 2023 · 被引用 11 次
- Mistify: Automating DNN Model Porting for On-Device Inference at the EdgePeizhen Guo, Bo Hu, Wenjun HuNSDI 2021 · 被引用 69 次
- Device-Wise Federated Network PruningShangqian Gao, Junyi Li, Zeyu Zhang, Yanfu Zhang 等CVPR 2024
- E4: Energy-Efficient DNN Inference for Edge Video Analytics via Early Exiting and DVFSZiyang Zhang, Yang Zhao, Ming-Ching Chang, Changyao Lin 等AAAI 2025 · 被引用 4 次
- Unlocking the Non-deterministic Computing Power with Memory-Elastic Multi-Exit Neural NetworksJiaming Huang, Yi Gao, Wei DongWWW 2024 · 被引用 3 次
