Re-thinking computation offload for efficient inference on IoT devices with duty-cycled radios
Jin Huang, Hui Guan, Deepak Ganesan
Abstract
While a number of recent efforts have explored the use of "cloud offload" to enable deep learning on IoT devices, these have not assumed the use of duty-cycled radios like BLE. We argue that radio duty-cycling significantly diminishes the performance of existing cloud-offload methods. We tackle this problem by leveraging a previously unexplored opportunity to use early-exit offload enhanced with prioritized communication, dynamic pooling, and dynamic fusion of features. We show that our system, FLEET, achieves significant benefits in accuracy, latency, and compute budget compared to state-of-art local early exit, remote processing, and model partitioning schemes across a range of DNN models, datasets, and IoT platforms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13de6cd3-8702-40d2-a6a3-aab4456c48dcCited by top-tier papers1
Ask how each one uses itBuilds on6
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis et al.MobiCom 2020 · 312 citations
- Memory-efficient Patch-based Inference for Tiny Deep LearningJi Lin, Wei-Ming Chen, Han Cai, Chuang Gan et al.NeurIPS 2021 · 190 citations
- Train Big, Then Compress: Rethinking Model Size for Efficient Training and Inference of TransformersZhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin et al.ICML 2020 · 184 citations
- CLIO: enabling automatic compilation of deep learning pipelines across IoT and cloudJin Huang, Colin Samplawski, Deepak Ganesan, Benjamin M. Marlin et al.MobiCom 2020 · 72 citations
Related papers
- Resource-aware Deployment of Dynamic DNNs over Multi-tiered Interconnected SystemsChetna Singhal, Yashuo Wu, Francesco Malandrino, Marco Levorato et al.INFOCOM 2024 · 15 citations
- Flex: Fast, Accurate DNN Inference on Low-Cost Edges Using Heterogeneous Accelerator ExecutionTanmoy Sen, Haiying Shen, Anand Padmanabha IyerEuroSys 2025 · 2 citations
- AccuMO: Accuracy-Centric Multitask Offloading in Edge-Assisted Mobile Augmented RealityZ. Jonny Kong, Qiang Xu, Jiayi Meng, Y. Charlie HuMobiCom 2023 · 21 citations
- Complement Sparsification: Low-Overhead Model Pruning for Federated LearningXiaopeng Jiang, Cristian BorceaAAAI 2023 · 36 citations
- Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient InferenceXiangjie Li, Chenfei Lou, Yuchi Chen, Zhengping Zhu et al.AAAI 2023 · 40 citations
