CLIO: enabling automatic compilation of deep learning pipelines across IoT and cloud
Jin Huang, Colin Samplawski, Deepak Ganesan, Benjamin M. Marlin, Heesung Kwon
Abstract
Recent years have seen dramatic advances in low-power neural accelerators that aim to bring deep learning analytics to IoT devices; simultaneously, there have been considerable advances in the design of low-power radios to enable efficient compute offload from IoT devices to the cloud. Neither is a panacea -deep learning models are often too large for low-power accelerators and bandwidth needs are often too high for low-power radios. While there has been considerable work on deep learning for smartphone-class devices, these methods do not work well for small battery-powered IoT devices that are considerably more resource-constrained.
In this paper, we bridge this gap by designing a continuously tunable method for leveraging both local and remote resources to optimize performance of a deep learning model. Clio presents a novel approach to split machine learning models between an IoT device and cloud in a progressive manner that adapts to wireless dynamics. We show that this method can be combined with model compression and adaptive model partitioning to create an integrated system for IoT-cloud partitioning. We implement Clio on the GAP8 low-power neural accelerator, provide an exhaustive characterization of the operating regimes where each method performs best and show that Clio can enable graceful performance degradation as resources diminish.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27020c1c-85b9-4988-b35d-eb9487a135b6Cited by top-tier papers7
- JAVP: Joint-Aware Video Processing with Edge-Cloud Collaboration for DNN InferenceZheming Yang, Wen Ji, Qi Guo, Zhi WangACM MM 2023 · 23 citations
- Re-thinking computation offload for efficient inference on IoT devices with duty-cycled radiosJin Huang, Hui Guan, Deepak GanesanMobiCom 2023 · 11 citations
- On the Robustness of Neural-Enhanced Video Streaming against Adversarial AttacksQihua Zhou, Jingcai Guo, Song Guo, Ruibin Li et al.AAAI 2024 · 7 citations
- Accelerating End-Cloud Collaborative Inference via Near Bubble-Free Pipeline OptimizationLuyao Gao, Jianchun Liu, Hongli Xu, Sun Xu et al.INFOCOM 2025 · 4 citations
- Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative LearningQihua Zhou, Song Guo, Yi Liu, Jie Zhang et al.NeurIPS 2022 · 4 citations
Builds on1
Related papers
- Progressive Neural Compression for Adaptive Image Offloading Under Timing ConstraintsRuiqi Wang, Hanyang Liu, Jiaming Qiu, Moran Xu et al.RTSS 2023 · 10 citations
- Context-Aware Compilation of DNN Training Pipelines across Edge and CloudDixi Yao, Liyao Xiang, Zifan Wang, Jiayu Xu et al.UbiComp 2022 · 25 citations
- Input-Dependent Edge-Cloud Mapping of Recurrent Neural Networks InferenceDaniele Jahier Pagliari, Roberta Chiaro, Yukai Chen, Sara Vinco et al.DAC 2020 · 9 citations
- Dynamic Resource Allocation for Deep Learning Clusters with Separated Compute and StorageMingxia Li, Zhenhua Han, Chi Zhang, Ruiting Zhou et al.INFOCOM 2023 · 3 citations
- Real-time neural network inference on extremely weak devices: agile offloading with explainable AIKai Huang, Wei GaoMobiCom 2022 · 57 citations
