FT-MoE: Sustainable-learning Mixture of Experts for Fault-Tolerant Computing
Wenjing Xiao, Wenhao Song, Miaojiang Chen, Min Chen
Abstract
Intelligent fault-tolerant (FT) computing has recently demonstrated significant advantages in predicting and diagnosing faults proactively, thereby ensuring reliable service delivery. However, due to the heterogeneity of fault knowledge, dynamic workloads, and limited data support, existing deep learning-based FT algorithms face challenges in fault detection quality and training efficiency. This is primarily because their homogenization of fault knowledge perception difficuties to fully capture diverse and complex fault patterns. To address these challenges, we propose FT-MoE, a sustainable-learning fault-tolerant computing framework based on a dual-path architecture for high-accuracy fault detection and classification. This model employs a mixture-of-experts (MoE) architecture, enabling different parameters to learn distinct fault knowledge. Additionally, we adopt a two-stage learning scheme that combines comprehensive offline training with continual online tuning, allowing the model to adaptively optimize its parameters in response to evolving real-time workloads. To facilitate realistic evaluation, we construct a new fault detection and classification dataset for edge networks, comprising 10,000 intervals with fine-grained resource features, surpassing existing datasets in both scale and granularity. Finally, we conduct extensive experiments on the FT benchmark to verify the effectiveness of FT-MoE. Results demonstrate that our model outperforms state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5705fc8-3fe8-4973-9c8a-d89aa8431be8Builds on5
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 1,717 citations
- TranAD: Deep Transformer Networks for Anomaly Detection in Multivariate Time Series DataShreshth Tuli, Giuliano Casale, Nicholas R. JenningsVLDB 2022 · 930 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
- MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model TrainingWeilin Cai, Le Qin, Jiayi HuangASPLOS 2025 · 5 citations
- Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of ExpertsXiaoming Shi, Shiyu Wang, Yuqi Nie, Dianqi Li et al.ICLR 2025
Related papers
- Theory of Mixture-of-Experts for Mobile Edge ComputingHongbo Li, Lingjie DuanINFOCOM 2025 · 9 citations
- DeepFT: Fault-Tolerant Edge Computing using a Self-Supervised Deep Surrogate ModelShreshth Tuli, Giuliano Casale, Ludmila Cherkasova, Nicholas R. JenningsINFOCOM 2023 · 23 citations
- FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device PlacementXiaonan Nie, Xupeng Miao, Zilong Wang, Zichao Yang et al.SIGMOD 2023 · 40 citations
- ADMoE: Anomaly Detection with Mixture-of-Experts from Noisy LabelsYue Zhao, Guoqing Zheng, Subhabrata Mukherjee, Robert McCann et al.AAAI 2023 · 39 citations
- Long-Tailed Visual Recognition via Self-Heterogeneous Integration with Knowledge ExcavationYan Jin, Mengke Li, Yang Lu, Yiu-ming Cheung et al.CVPR 2023
