CLIO: enabling automatic compilation of deep learning pipelines across IoT and cloud
Jin Huang, Colin Samplawski, Deepak Ganesan, Benjamin M. Marlin, Heesung Kwon
摘要
Recent years have seen dramatic advances in low-power neural accelerators that aim to bring deep learning analytics to IoT devices; simultaneously, there have been considerable advances in the design of low-power radios to enable efficient compute offload from IoT devices to the cloud. Neither is a panacea -deep learning models are often too large for low-power accelerators and bandwidth needs are often too high for low-power radios. While there has been considerable work on deep learning for smartphone-class devices, these methods do not work well for small battery-powered IoT devices that are considerably more resource-constrained.
In this paper, we bridge this gap by designing a continuously tunable method for leveraging both local and remote resources to optimize performance of a deep learning model. Clio presents a novel approach to split machine learning models between an IoT device and cloud in a progressive manner that adapts to wireless dynamics. We show that this method can be combined with model compression and adaptive model partitioning to create an integrated system for IoT-cloud partitioning. We implement Clio on the GAP8 low-power neural accelerator, provide an exhaustive characterization of the operating regimes where each method performs best and show that Clio can enable graceful performance degradation as resources diminish.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- JAVP: Joint-Aware Video Processing with Edge-Cloud Collaboration for DNN InferenceZheming Yang, Wen Ji, Qi Guo, Zhi WangACM MM 2023 · 被引用 23 次
- Re-thinking computation offload for efficient inference on IoT devices with duty-cycled radiosJin Huang, Hui Guan, Deepak GanesanMobiCom 2023 · 被引用 11 次
- On the Robustness of Neural-Enhanced Video Streaming against Adversarial AttacksQihua Zhou, Jingcai Guo, Song Guo, Ruibin Li 等AAAI 2024 · 被引用 7 次
- Accelerating End-Cloud Collaborative Inference via Near Bubble-Free Pipeline OptimizationLuyao Gao, Jianchun Liu, Hongli Xu, Sun Xu 等INFOCOM 2025 · 被引用 4 次
- Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative LearningQihua Zhou, Song Guo, Yi Liu, Jie Zhang 等NeurIPS 2022 · 被引用 4 次
它引用的顶会 Paper1
相关 Paper
- Progressive Neural Compression for Adaptive Image Offloading Under Timing ConstraintsRuiqi Wang, Hanyang Liu, Jiaming Qiu, Moran Xu 等RTSS 2023 · 被引用 10 次
- Context-Aware Compilation of DNN Training Pipelines across Edge and CloudDixi Yao, Liyao Xiang, Zifan Wang, Jiayu Xu 等UbiComp 2022 · 被引用 25 次
- Input-Dependent Edge-Cloud Mapping of Recurrent Neural Networks InferenceDaniele Jahier Pagliari, Roberta Chiaro, Yukai Chen, Sara Vinco 等DAC 2020 · 被引用 9 次
- Dynamic Resource Allocation for Deep Learning Clusters with Separated Compute and StorageMingxia Li, Zhenhua Han, Chi Zhang, Ruiting Zhou 等INFOCOM 2023 · 被引用 3 次
- Real-time neural network inference on extremely weak devices: agile offloading with explainable AIKai Huang, Wei GaoMobiCom 2022 · 被引用 57 次
