Hierarchical Inference with an Offload Queue
Srinivas Nomula, Parimal Parag, Ayalvadi Ganesh, Jaya Prakash Champati
Abstract
In a Hierarchical Inference (HI) system, an end device processes data samples using a local Deep Learning (DL) model. Only when the confidence of the local DL is less than a chosen threshold, it offloads the sample for remote DL inference on an edge server. By balancing the misclassifications and offloading costs, a HI system can improve accuracy and responsiveness and reduce bandwidth usage. To this end, prior work studied a HI learning problem for classification tasks, aiming to learn an optimal threshold on the confidence to minimize the cumulative costs. However, existing formulations considered (i) abstract offloading costs and (ii) assumed that the remote DL inference for an offloaded task is available by the end of each round. In contrast, considering that end devices are wirelessly connected to edge servers, we study a HI system with an offload queue where the offloading cost is implicitly modeled by the queue length. Also, we consider the practical setting, where the remote DL inference is obtained only after the task is served from the offload queue. For bounded but unknown arrival and service processes, we formulate the problem of average mismatch error minimization subject to queue stability. We propose a BOLD Hedge-Q policy based on Lyapunov optimization. The novelty of BOLD Hedge-Q lies in solving a delayed-feedback online learning problem — with sub-linear regret — that results from the Lyapunov drift-plus-penalty analysis. Given a penalty parameter V ≥ 1, we show that under BOLD Hedge-Q, the queue is upper bounded by V + λm, and the mismatch error is within an additive factor of O(1/V) from that of the optimal policy, where λm is the maximum number of arrivals in a round.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on2
- Distributed Inference with Deep Learning Models across Heterogeneous Edge DevicesChenghao Hu, Baochun LiINFOCOM 2022 · 81 citations
- Optimization of Offloading Policies for Accuracy-Delay Tradeoffs in Hierarchical InferenceHasan Burhan Beytur, Ahmet Günhan Aydin, Gustavo de Veciana, Haris VikaloINFOCOM 2024 · 11 citations
Related papers
- Inference Offloading for Cost-Sensitive Binary Classification at the EdgeVishnu Narayanan Moothedath, Umang Agarwal, Umeshraja N, James Gross et al.AAAI 2026
- Autodidactic Neurosurgeon: Collaborative Deep Inference for Mobile Edge Intelligence via Online LearningLetian Zhang, Lixing Chen, Jie XuWWW 2021 · 75 citations
- Decentralized Task Offloading in Edge Computing: A Multi-User Multi-Armed Bandit ApproachXiong Wang, Jiancheng Ye, John C. S. LuiINFOCOM 2022 · 89 citations
- Edge-assisted Online On-device Object Detection for Real-time Video AnalyticsMengxi Hanyao, Yibo Jin, Zhuzhong Qian, Sheng Zhang et al.INFOCOM 2021 · 82 citations
- Joint Optimization of Model Deployment for Freshness-Sensitive Task Assignment in Edge IntelligenceHaolin Liu, Sirui Liu, Saiqin Long, Qingyong Deng et al.INFOCOM 2024 · 14 citations
