Minder: Faulty Machine Detection for Large-scale Distributed Model Training
Yangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang, Lei Zhang, Zhang Zhang, Bo Li, Zuquan Song, Hang Zhu, Gaohong Liu, Fuliang Li, Shuguang Wang
摘要
Large-scale distributed model training requires simultaneous training on up to thousands of machines. Faulty machine detection is critical when an unexpected fault occurs in a machine. From our experience, a training task can encounter two faults per day on average, possibly leading to a halt for hours. To address the drawbacks of the time-consuming and labor-intensive manual scrutiny, we propose Minder, an automatic faulty machine detector for distributed training tasks. The key idea of Minder is to automatically and efficiently detect faulty distinctive monitoring metric patterns, which could last for a period before the entire training task comes to a halt. Minder has been deployed in our production environment for over one year, monitoring daily distributed training tasks where each involves up to thousands of machines. In our real-world fault detection scenarios, Minder can accurately and efficiently react to faults within 3.6 seconds on average, with a precision of 0.904 and F1-score of 0.893.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Astral: A Datacenter Infrastructure for Large Language Model Training at ScaleQingkai Meng, Hao Zheng, Zhenhui Zhang, ChonLam Lao 等SIGCOMM 2025 · 被引用 16 次
- Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM TrainingYangtao Deng, Lei Zhang, Qinlong Wang, Xiaoyun Zhi 等SOSP 2025 · 被引用 7 次
- CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingYida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun 等PPoPP 2026
- RobustRL: Role-Based Fault Tolerance System for RL Post-TrainingZhenqian Chen, Baoquan Zhong, Xiang Li, Qing Dai 等OSDI 2026
- TSGuard: Automated User-Centric Incident Diagnosis for AI Workloads in the CloudYitao Yang, Yangtao Deng, Yifan Xiong, Baochun Li 等FSE 2026
它引用的顶会 Paper15
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 等NSDI 2024 · 被引用 415 次
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang 等NSDI 2024 · 被引用 192 次
- Sage: practical and scalable ML-driven performance debugging in microservicesYu Gan, Mingyu Liang, Sundar Dev, David Lo 等ASPLOS 2021 · 被引用 170 次
- Interpreting Deep Learning-Based Networking SystemsZili Meng, Minhu Wang, Jiasong Bai, Mingwei Xu 等SIGCOMM 2020 · 被引用 98 次
- Collie: Finding Performance Anomalies in RDMA SubsystemsXinhao Kong, Yibo Zhu, Huaping Zhou, Zhuo Jiang 等NSDI 2022 · 被引用 86 次
相关 Paper
- Evolution of Aegis: Fault Diagnosis for AI Model Training Service in ProductionJianbo Dong, Kun Qian, Pengcheng Zhang, Zhilong Zheng 等NSDI 2025 · 被引用 21 次
- GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang 等USENIX ATC 2025 · 被引用 19 次
- EROICA: Online Performance Troubleshooting for Large-scale Model TrainingYu Guan, Zhiyu Yin, Haoyu Chen, Sheng Cheng 等NSDI 2026 · 被引用 1 次
- TrainMover: An Interruption-Resilient Runtime for ML TrainingChonLam Lao, Jiaqi Gao, Jiamin Cao, Zhipeng Zhang 等OSDI 2026
- PCcheck: Persistent Concurrent Checkpointing for MLFoteini Strati, Michal Friedman, Ana KlimovicASPLOS 2025 · 被引用 11 次
