Real-time neural network inference on extremely weak devices: agile offloading with explainable AI
Kai Huang, Wei Gao
摘要
With the wide adoption of AI applications, there is a pressing need of enabling real-time neural network (NN) inference on small embedded devices, but deploying NNs and achieving high performance of NN inference on these small devices is challenging due to their extremely weak capabilities. Although NN partitioning and offloading can contribute to such deployment, they are incapable of minimizing the local costs at embedded devices. Instead, we suggest to address this challenge via agile NN offloading, which migrates the required computations in NN offloading from online inference to offline learning. In this paper, we present AgileNN, a new NN offloading technique that achieves real-time NN inference on weak embedded devices by leveraging eXplainable AI techniques, so as to explicitly enforce feature sparsity during the training phase and minimize the online computation and communication costs. Experiment results show that AgileNN's inference latency is >6X lower than the existing schemes, ensuring that sensory data on embedded devices can be timely consumed. It also reduces the local device's resource consumption by >8X, without impairing the inference accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge DevicesShengyuan Ye, Liekang Zeng, Xiaowen Chu, Guoliang Xing 等MobiCom 2024 · 被引用 29 次
- DTMM: Deploying TinyML Models on Extremely Weak IoT Devices with PruningLixiang Han, Zhen Xiao, Zhenjiang LiINFOCOM 2024 · 被引用 20 次
- AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT ApplicationsLeming Shen, Qiang Yang, Yuanqing Zheng, Mo LiMobiCom 2025 · 被引用 14 次
- LATTE: Layer Algorithm-aware Training Time Estimation for Heterogeneous Federated LearningKun Wang, Zimu Zhou, Zhenjiang LiMobiCom 2024 · 被引用 13 次
- TensorShield: Safeguarding On-Device Inference by Shielding Critical DNN Tensors with TEETong Sun, Bowen Jiang, Hailong Lin, Borui Li 等CCS 2025 · 被引用 3 次
它引用的顶会 Paper10
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 被引用 4,239 次
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn 等NeurIPS 2020 · 被引用 827 次
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis 等MobiCom 2020 · 被引用 312 次
- Pruning Neural Networks at Initialization: Why Are We Missing the Mark?Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICLR 2021 · 被引用 261 次
- Deep learning based wireless localization for indoor navigationRoshan Sai Ayyalasomayajula, Aditya Arun, Chenfeng Wu, Sanatan Sharma 等MobiCom 2020 · 被引用 221 次
相关 Paper
- AoDNN: An Auto-Offloading Approach to Optimize Deep Inference for Fostering Mobile WebYakun Huang, Xiuquan Qiao, Schahram Dustdar, Yan LiINFOCOM 2022 · 被引用 17 次
- SemanticNN: Compressive and Error-Resilient Semantic Offloading for Extremely Weak DevicesJiaming Huang, Yi Gao, Fuchang Pan, Renjie Li 等AAAI 2026
- Distributed Inference Acceleration with Adaptive DNN Partitioning and OffloadingThaha Mohammed, Carlee Joe-Wong, Rohit Babbar, Mario Di FrancescoINFOCOM 2020 · 被引用 213 次
- CoExe: An Efficient Co-execution Architecture for Real-Time Neural Network ServicesChubo Liu, Kenli Li, Mingcong Song, Jiechen Zhao 等DAC 2020 · 被引用 4 次
- CLIO: enabling automatic compilation of deep learning pipelines across IoT and cloudJin Huang, Colin Samplawski, Deepak Ganesan, Benjamin M. Marlin 等MobiCom 2020 · 被引用 72 次
