Evolution of Aegis: Fault Diagnosis for AI Model Training Service in Production
Jianbo Dong, Kun Qian, Pengcheng Zhang, Zhilong Zheng, Liang Chen, Fei Feng, Yichi Xu, Yikai Zhu, Gang Lu, Xue Li, Zhihui Ren, Zhicheng Wang
摘要
Despite the success of diagnosis systems in traditional cloud computing, these systems are not suitable for pinpointing faults in AI model training cloud scenarios due to the differences in computing paradigms between traditional cloud computing and model training. As one of the largest cloud providers, we present Aegis, a fault diagnosis system specifically designed for AI model training service. We share our experience in the motivation, design, and evolution of Aegis. Keeping easy-to-deploy as the primary principle, Aegis Phase-1 started by enhancing existing general-purpose diagnosis systems. After several months of evolution, Aegis Phase-2 cogitatively customized the collective communication library for sophisticated failure localization in runtime without modifying customer code. Besides the failure localization, we further equipped Aegis with the capabilities on handling performance degradation and failure checking before delivery. Aegis has been deployed in our production training cloud service for one year. Aegis decreases more than 97% of the idle time wasted by diagnosis, 84% of the training task restart count, and 71% of the performance degradation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM TrainingYangtao Deng, Lei Zhang, Qinlong Wang, Xiaoyun Zhi 等SOSP 2025 · 被引用 7 次
- Laminar: A Scalable Asynchronous RL Post-Training FrameworkGuangming Sheng, Yuxuan Tong, Borui Wan, Wang Zhang 等EuroSys 2026 · 被引用 2 次
- EROICA: Online Performance Troubleshooting for Large-scale Model TrainingYu Guan, Zhiyu Yin, Haoyu Chen, Sheng Cheng 等NSDI 2026 · 被引用 1 次
- Robust LLM Training Infrastructure at ByteDanceBorui Wan, Gaohong Liu, Zuquan Song, Jun Wang 等SOSP 2025 · 被引用 1 次
- TSGuard: Automated User-Centric Incident Diagnosis for AI Workloads in the CloudYitao Yang, Yangtao Deng, Yifan Xiong, Baochun Li 等FSE 2026
它引用的顶会 Paper18
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 等NSDI 2024 · 被引用 415 次
- PINT: Probabilistic In-band Network TelemetryRan Ben Basat, Sivaramakrishnan Ramanathan, Yuliang Li, Gianni Antichi 等SIGCOMM 2020 · 被引用 268 次
- Alibaba HPN: A Data Center Network for Large Language Model TrainingKun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao 等SIGCOMM 2024 · 被引用 173 次
- Flow Event Telemetry on Programmable Data PlaneYu Zhou, Chen Sun, Hongqiang Harry Liu, Rui Miao 等SIGCOMM 2020 · 被引用 139 次
- Collie: Finding Performance Anomalies in RDMA SubsystemsXinhao Kong, Yibo Zhu, Huaping Zhou, Zhuo Jiang 等NSDI 2022 · 被引用 86 次
相关 Paper
- Safeguarding LLM Training at Scale: Online SDC Detection and Insights from 35 Million GPU HoursKinman Lei, Liyan Zheng, Xiang Li, Hongmin Chen 等OSDI 2026
- Minder: Faulty Machine Detection for Large-scale Distributed Model TrainingYangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang 等NSDI 2025 · 被引用 36 次
- Handling Network Faults in Distributed AI Training: Failover is Now an OptionXin Zhe Khooi, Zhuo Jiang, Pan Xie, Zhigang Cui 等EuroSys 2026 · 被引用 1 次
- FLARE: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus ScaleWeihao Cui, Ji Zhang, Han Zhao, Chao Liu 等NSDI 2026 · 被引用 7 次
- Aegis: Contract-Bounded Online Adaptation for Networked Accelerator ClustersRui Li, Shuang CaoSIGCOMM 2026
