Understanding and Mitigating Hardware Failures in Deep Learning Training Systems
Yi He, Mike Hutton, Steven Chan, Robert De Gruijl, Rama Govindaraju, Nishant Patil, Yanjing Li
2023Year
52Citations
7Top-tier citations
Abstract
Deep neural network (DNN) training workloads are increasingly susceptible to hardware failures in datacenters. For example, Google experienced "mysterious, difficult to identify problems" in their TPU training systems due to hardware failures [7]. Although these particular problems were subsequently corrected through significant efforts, they have raised the urgency of addressing the growing challenges emerging from hardware failures impacting many DNN training workloads.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 26c6d420-b8a0-42b3-a1f6-2072c6aaf632Cited by top-tier papers7
- Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep LearningWei An, Xiao Bi, Guanting Chen, Shanhuang Chen et al.SC 2024 · 20 citations
- PCcheck: Persistent Concurrent Checkpointing for MLFoteini Strati, Michal Friedman, Ana KlimovicASPLOS 2025 · 11 citations
- ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model TrainingYuhang Liang, Xinyi Li, Jie Ren, Ang Li et al.PPoPP 2025 · 10 citations
- Proactive Runtime Detection of Aging-Related Silent Data Corruptions: A Bottom-Up ApproachJiacheng Ma, Majd Ganaiem, Madeline Burbage, Theo Gregersen et al.ASPLOS 2024 · 7 citations
- Robust LLM Training Infrastructure at ByteDanceBorui Wan, Gaohong Liu, Zuquan Song, Jun Wang et al.SOSP 2025 · 1 citation
Related papers
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev et al.EuroSys 2024 · 23 citations
- Understanding Stragglers in Large Model Training Using What-if AnalysisJinkun Lin, Ziheng Jiang, Zuquan Song, Sida Zhao et al.OSDI 2025 · 23 citations
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen et al.NSDI 2021 · 359 citations
- Resiliency at Scale: Managing Google's TPUv4 Machine Learning SupercomputerYazhou Zu, Alireza Ghaffarkhah, Hoang-Vu Dang, Brian Towles et al.NSDI 2024 · 46 citations
- Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated SystemsYonatan Levitt, Richard Barella, Sam Zeltner, Thomas Musta et al.SC 2025 · 1 citation
