SIEVE: Speculative Inference on the Edge with Versatile Exportation
Babak Zamirai, Salar Latifi, Pedram Zamirai, Scott A. Mahlke
摘要
This paper proposes SIEVE, Speculative Inference on the Edge with Versatile Exportation, which dynamically distributes CNN computation between the cloud and edge device based on the input data and environmental conditions to maximize efficiency and performance. A speculative CNN is created through aggressive precision reduction techniques to run most of the inferences on the edge device, while the original CNN is run on the cloud server. A runtime system directs each input to either the edge or cloud and decides whether to accept speculative inferences made on the edge or invoke recovery by replaying the inference on the cloud. Compared to the cloud-only approach, SIEVE reduces energy consumption by an average of 91%, 57% and 26% and increases performance by an average of 12.3×, 2.8× and 2.0× for 3G, LTE and WiFi connections without accuracy loss across a range of nine CNNs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LENS: Layer Distribution Enabled Neural Architecture Search in Edge-Cloud HierarchiesMohanad Odema, Nafiul Rashid, Berken Utku Demirel, Mohammad Abdullah Al FaruqueDAC 2021 · 被引用 8 次
- SEO: Safety-Aware Energy Optimization Framework for Multi-Sensor Neural Controllers at the EdgeMohanad Odema, James Ferlez, Yasser Shoukry, Mohammad Abdullah Al FaruqueDAC 2023 · 被引用 1 次
相关 Paper
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis 等MobiCom 2020 · 被引用 312 次
- Input-Dependent Edge-Cloud Mapping of Recurrent Neural Networks InferenceDaniele Jahier Pagliari, Roberta Chiaro, Yukai Chen, Sara Vinco 等DAC 2020 · 被引用 9 次
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia 等MobiCom 2021 · 被引用 171 次
- AppealNet: An Efficient and Highly-Accurate Edge/Cloud Collaborative Architecture for DNN InferenceMin Li, Yu Li, Ye Tian, Li Jiang 等DAC 2021 · 被引用 37 次
- QoS-Aware Irregular Collaborative Inference for Improving Throughput of DNN ServicesKaihua Fu, Jiuchen Shi, Quan Chen, Ningxin Zheng 等SC 2022 · 被引用 7 次
