SIEVE: Speculative Inference on the Edge with Versatile Exportation
Babak Zamirai, Salar Latifi, Pedram Zamirai, Scott A. Mahlke
Abstract
This paper proposes SIEVE, Speculative Inference on the Edge with Versatile Exportation, which dynamically distributes CNN computation between the cloud and edge device based on the input data and environmental conditions to maximize efficiency and performance. A speculative CNN is created through aggressive precision reduction techniques to run most of the inferences on the edge device, while the original CNN is run on the cloud server. A runtime system directs each input to either the edge or cloud and decides whether to accept speculative inferences made on the edge or invoke recovery by replaying the inference on the cloud. Compared to the cloud-only approach, SIEVE reduces energy consumption by an average of 91%, 57% and 26% and increases performance by an average of 12.3×, 2.8× and 2.0× for 3G, LTE and WiFi connections without accuracy loss across a range of nine CNNs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a7137b2-ec44-46d9-8245-b1b26c22ff5eCited by top-tier papers2
- LENS: Layer Distribution Enabled Neural Architecture Search in Edge-Cloud HierarchiesMohanad Odema, Nafiul Rashid, Berken Utku Demirel, Mohammad Abdullah Al FaruqueDAC 2021 · 8 citations
- SEO: Safety-Aware Energy Optimization Framework for Multi-Sensor Neural Controllers at the EdgeMohanad Odema, James Ferlez, Yasser Shoukry, Mohammad Abdullah Al FaruqueDAC 2023 · 1 citation
Related papers
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis et al.MobiCom 2020 · 312 citations
- Input-Dependent Edge-Cloud Mapping of Recurrent Neural Networks InferenceDaniele Jahier Pagliari, Roberta Chiaro, Yukai Chen, Sara Vinco et al.DAC 2020 · 9 citations
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia et al.MobiCom 2021 · 171 citations
- AppealNet: An Efficient and Highly-Accurate Edge/Cloud Collaborative Architecture for DNN InferenceMin Li, Yu Li, Ye Tian, Li Jiang et al.DAC 2021 · 37 citations
- QoS-Aware Irregular Collaborative Inference for Improving Throughput of DNN ServicesKaihua Fu, Jiuchen Shi, Quan Chen, Ningxin Zheng et al.SC 2022 · 7 citations
