DeepFT: Fault-Tolerant Edge Computing using a Self-Supervised Deep Surrogate Model
Shreshth Tuli, Giuliano Casale, Ludmila Cherkasova, Nicholas R. Jennings
摘要
The emergence of latency-critical AI applications has been supported by the evolution of the edge computing paradigm. However, edge solutions are typically resourceconstrained, posing reliability challenges due to heightened contention for compute and communication capacities and faulty application behavior in the presence of overload conditions. Although a large amount of generated log data can be mined for fault prediction, labeling this data for training is a manual process and thus a limiting factor for automation. Due to this, many companies resort to unsupervised fault-tolerance models. Yet, failure models of this kind can incur a loss of accuracy when they need to adapt to non-stationary workloads and diverse host characteristics. To cope with this, we propose a novel modeling approach, called DeepFT, to proactively avoid system overloads and their adverse effects by optimizing the task scheduling and migration decisions. DeepFT uses a deep surrogate model to accurately predict and diagnose faults in the system and cosimulation based self-supervised learning to dynamically adapt the model in volatile settings. It offers a highly scalable solution as the model size scales by only 3 and 1 percent per unit increase in the number of active tasks and hosts. Extensive experimentation on a Raspberry-Pi based edge cluster with DeFog benchmarks shows that DeepFT can outperform state-of-the-art baseline methods in fault-detection and QoS metrics. Specifically, DeepFT gives the highest F1 scores for fault-detection, reducing service deadline violations by up to 37% while also improving response time by up to 9%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- FT-MoE: Sustainable-learning Mixture of Experts for Fault-Tolerant ComputingWenjing Xiao, Wenhao Song, Miaojiang Chen, Min ChenAAAI 2026 · 被引用 1 次
- AdaInf: Data Drift Adaptive Scheduling for Accurate and SLO-guaranteed Multiple-Model Inference Serving at Edge ServersSudipta Saha Shubha, Haiying ShenSIGCOMM 2023 · 被引用 34 次
- RESCUE: Opportunistic Online Scheduling of Model Retraining on Underutilized EdgesJianping Huang, Xiang Liu, Feng ShanINFOCOM 2026
- Dissector: input validation for deep learning applications by crossing-layer dissectionHuiyan Wang, Jingwei Xu, Chang Xu, Xiaoxing Ma 等ICSE 2020 · 被引用 56 次
- Intelligent Resource Scheduling for Co-located Latency-critical Services: A Multi-Model Collaborative Learning ApproachLei Liu, Xinglei Dou, Yuetao ChenFAST 2023
