Inference Offloading for Cost-Sensitive Binary Classification at the Edge
Vishnu Narayanan Moothedath, Umang Agarwal, Umeshraja N, James Gross, Jaya Prakash Champati, Sharayu Moharir
摘要
We investigate a binary classification problem in an edge intelligence system where false negatives are more costly than false positives. The system features a compact, locally deployed model, supplemented by a larger, remote model that is accessible via the network, albeit at an offloading cost. For each sample, our system first uses the locally deployed model for inference. Based on the output of the local model, the sample may be offloaded to the remote model. This work aims to understand the fundamental trade-off between classification accuracy and the offloading costs within such a hierarchical inference (HI) system. To optimise this system, we propose an online learning framework that continuously adapts a pair of thresholds on the local model's confidence scores. These thresholds determine the prediction of the local model and whether a sample is classified locally or offloaded to the remote model. We present a closed-form solution for the setting where the local model is calibrated. For the more general case of uncalibrated models, we introduce H2T2, an online two-threshold hierarchical inference policy, and prove it achieves sublinear regret. H2T2 is model-agnostic, requires no training, and learns during the inference phase using limited feedback. Simulations on real-world datasets show that H2T2 consistently outperforms naive and single-threshold HI policies, sometimes even surpassing single-threshold offline optima. The policy also demonstrates robustness to distribution shifts and adapts effectively to mismatched classifiers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- Post-hoc estimators for learning to defer to an expertHarikrishna Narasimhan, Wittawat Jitkrittum, Aditya Krishna Menon, Ankit Singh Rawat 等NeurIPS 2022 · 被引用 66 次
- Optimization of Offloading Policies for Accuracy-Delay Tradeoffs in Hierarchical InferenceHasan Burhan Beytur, Ahmet Günhan Aydin, Gustavo de Veciana, Haris VikaloINFOCOM 2024 · 被引用 11 次
相关 Paper
- Hierarchical Inference with an Offload QueueSrinivas Nomula, Parimal Parag, Ayalvadi Ganesh, Jaya Prakash ChampatiINFOCOM 2026
- Edge-assisted Online On-device Object Detection for Real-time Video AnalyticsMengxi Hanyao, Yibo Jin, Zhuzhong Qian, Sheng Zhang 等INFOCOM 2021 · 被引用 82 次
- AI in 5G: The Case of Online Distributed Transfer Learning over Edge NetworksYulan Yuan, Lei Jiao, Konglin Zhu, Xiaojun Lin 等INFOCOM 2022 · 被引用 13 次
- Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and InferenceHuaiguang Cai, Zhi Zhou, Qianyi HuangINFOCOM 2024 · 被引用 10 次
- Online Learning with Primary and Secondary LossesAvrim Blum, Han ShaoNeurIPS 2020 · 被引用 1 次
