Jellyfish: Timely Inference Serving for Dynamic Edge Networks
Vinod Nigade, Pablo Bauszat, Henri E. Bal, Lin Wang
摘要
While high accuracy is of paramount importance for deep learning (DL) inference, serving inference requests on time is equally critical but has not been carefully studied especially when the request has to be served over a dynamic wireless network at the edge. In this paper, we propose Jellyfish—a novel edge DL inference serving system that achieves soft guarantees on end-to-end inference latency often specified as a service-level objective (SLO). To handle the network variability, Jellyfish exploits both data and deep neural network (DNN) adaptation to conduct tradeoffs between accuracy and latency. Jellyfish features a new design that enables collective adaptation policies where the decisions for data and DNN adaptations are aligned and coordinated among multiple users with varying network conditions. We propose efficient algorithms to dynamically adapt DNNs and map users, so that we fulfill latency SLOs while maximizing the overall inference accuracy. Our experiments based on a prototype implementation and real-world WiFi and LTE network traces show that Jellyfish can meet latency SLOs at around the 99th percentile while maintaining high accuracy.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Model Selection for Latency-Critical Inference ServingDaniel Mendoza, Francisco Romero, Caroline TrippelEuroSys 2024 · 被引用 16 次
- RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language ModelsYufei Li, Zexin Li, Wei Yang, Cong LiuRTSS 2023 · 被引用 10 次
- : On-Device Real-Time Deep Reinforcement Learning for Autonomous RoboticsZexin Li, Aritra Samanta, Yufei Li, Andrea Soltoggio 等RTSS 2023 · 被引用 9 次
- X-Stream: A Flexible, Adaptive Video Transformer for Privacy-Preserving Video Stream AnalyticsDou Feng, Lin Wang, Shutong Chen, Lingching Tung 等INFOCOM 2024 · 被引用 8 次
- Uirapuru: Timely Video Analytics for High-Resolution Steerable Cameras on Edge DevicesGuilherme Henrique Apostolo, Pablo Bauszat, Vinod Nigade, Henri E. Bal 等MobiCom 2025 · 被引用 1 次
相关 Paper
- AdaInf: Data Drift Adaptive Scheduling for Accurate and SLO-guaranteed Multiple-Model Inference Serving at Edge ServersSudipta Saha Shubha, Haiying ShenSIGCOMM 2023 · 被引用 34 次
- DeepAdapter: A Collaborative Deep Learning Framework for the Mobile Web Using Context-Aware Network PruningYakun Huang, Xiuquan Qiao, Jian Tang, Pei Ren 等INFOCOM 2020 · 被引用 32 次
- Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AIZhi Zhou, Jiajie Xie, Mengke Huang, Tao Ouyang 等INFOCOM 2025 · 被引用 3 次
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao 等OSDI 2020 · 被引用 392 次
- ALERT: Accurate Learning for Energy and TimelinessChengcheng Wan, Muhammad Husni Santriaji, Eri Rogers, Henry Hoffmann 等USENIX ATC 2020 · 被引用 15 次
