Jellyfish: Timely Inference Serving for Dynamic Edge Networks
Vinod Nigade, Pablo Bauszat, Henri E. Bal, Lin Wang
Abstract
While high accuracy is of paramount importance for deep learning (DL) inference, serving inference requests on time is equally critical but has not been carefully studied especially when the request has to be served over a dynamic wireless network at the edge. In this paper, we propose Jellyfish—a novel edge DL inference serving system that achieves soft guarantees on end-to-end inference latency often specified as a service-level objective (SLO). To handle the network variability, Jellyfish exploits both data and deep neural network (DNN) adaptation to conduct tradeoffs between accuracy and latency. Jellyfish features a new design that enables collective adaptation policies where the decisions for data and DNN adaptations are aligned and coordinated among multiple users with varying network conditions. We propose efficient algorithms to dynamically adapt DNNs and map users, so that we fulfill latency SLOs while maximizing the overall inference accuracy. Our experiments based on a prototype implementation and real-world WiFi and LTE network traces show that Jellyfish can meet latency SLOs at around the 99th percentile while maintaining high accuracy.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 39ec803a-d6ea-49f6-a8e1-bfe244dcf994Cited by top-tier papers5
- Model Selection for Latency-Critical Inference ServingDaniel Mendoza, Francisco Romero, Caroline TrippelEuroSys 2024 · 16 citations
- RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language ModelsYufei Li, Zexin Li, Wei Yang, Cong LiuRTSS 2023 · 10 citations
- : On-Device Real-Time Deep Reinforcement Learning for Autonomous RoboticsZexin Li, Aritra Samanta, Yufei Li, Andrea Soltoggio et al.RTSS 2023 · 9 citations
- X-Stream: A Flexible, Adaptive Video Transformer for Privacy-Preserving Video Stream AnalyticsDou Feng, Lin Wang, Shutong Chen, Lingching Tung et al.INFOCOM 2024 · 8 citations
- Uirapuru: Timely Video Analytics for High-Resolution Steerable Cameras on Edge DevicesGuilherme Henrique Apostolo, Pablo Bauszat, Vinod Nigade, Henri E. Bal et al.MobiCom 2025 · 1 citation
Related papers
- AdaInf: Data Drift Adaptive Scheduling for Accurate and SLO-guaranteed Multiple-Model Inference Serving at Edge ServersSudipta Saha Shubha, Haiying ShenSIGCOMM 2023 · 34 citations
- DeepAdapter: A Collaborative Deep Learning Framework for the Mobile Web Using Context-Aware Network PruningYakun Huang, Xiuquan Qiao, Jian Tang, Pei Ren et al.INFOCOM 2020 · 32 citations
- Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AIZhi Zhou, Jiajie Xie, Mengke Huang, Tao Ouyang et al.INFOCOM 2025 · 3 citations
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- ALERT: Accurate Learning for Energy and TimelinessChengcheng Wan, Muhammad Husni Santriaji, Eri Rogers, Henry Hoffmann et al.USENIX ATC 2020 · 15 citations
