FT-MoE: Sustainable-learning Mixture of Experts for Fault-Tolerant Computing
Wenjing Xiao, Wenhao Song, Miaojiang Chen, Min Chen
摘要
Intelligent fault-tolerant (FT) computing has recently demonstrated significant advantages in predicting and diagnosing faults proactively, thereby ensuring reliable service delivery. However, due to the heterogeneity of fault knowledge, dynamic workloads, and limited data support, existing deep learning-based FT algorithms face challenges in fault detection quality and training efficiency. This is primarily because their homogenization of fault knowledge perception difficuties to fully capture diverse and complex fault patterns. To address these challenges, we propose FT-MoE, a sustainable-learning fault-tolerant computing framework based on a dual-path architecture for high-accuracy fault detection and classification. This model employs a mixture-of-experts (MoE) architecture, enabling different parameters to learn distinct fault knowledge. Additionally, we adopt a two-stage learning scheme that combines comprehensive offline training with continual online tuning, allowing the model to adaptively optimize its parameters in response to evolving real-time workloads. To facilitate realistic evaluation, we construct a new fault detection and classification dataset for edge networks, comprising 10,000 intervals with fine-grained resource features, surpassing existing datasets in both scale and granularity. Finally, we conduct extensive experiments on the FT benchmark to verify the effectiveness of FT-MoE. Results demonstrate that our model outperforms state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 被引用 1,717 次
- TranAD: Deep Transformer Networks for Anomaly Detection in Multivariate Time Series DataShreshth Tuli, Giuliano Casale, Nicholas R. JenningsVLDB 2022 · 被引用 930 次
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu 等ACL 2024 · 被引用 171 次
- MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model TrainingWeilin Cai, Le Qin, Jiayi HuangASPLOS 2025 · 被引用 5 次
- Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of ExpertsXiaoming Shi, Shiyu Wang, Yuqi Nie, Dianqi Li 等ICLR 2025
相关 Paper
- Theory of Mixture-of-Experts for Mobile Edge ComputingHongbo Li, Lingjie DuanINFOCOM 2025 · 被引用 9 次
- DeepFT: Fault-Tolerant Edge Computing using a Self-Supervised Deep Surrogate ModelShreshth Tuli, Giuliano Casale, Ludmila Cherkasova, Nicholas R. JenningsINFOCOM 2023 · 被引用 23 次
- FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device PlacementXiaonan Nie, Xupeng Miao, Zilong Wang, Zichao Yang 等SIGMOD 2023 · 被引用 40 次
- ADMoE: Anomaly Detection with Mixture-of-Experts from Noisy LabelsYue Zhao, Guoqing Zheng, Subhabrata Mukherjee, Robert McCann 等AAAI 2023 · 被引用 39 次
- Long-Tailed Visual Recognition via Self-Heterogeneous Integration with Knowledge ExcavationYan Jin, Mengke Li, Yang Lu, Yiu-ming Cheung 等CVPR 2023
