SC2025Top-tier venue
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
Shengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan, Ziheng Chen, Phuong Cao, Gregory H. Bauer, Brett M. Bode, Catello Di Martino, Saurabh Jha, Chandra Narayanaswami, Daby Sow
Abstract
This study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years of operational data (11.7 million GPU hours) on GPU errors. Our major findings include: (i) H100 GPU memory resilience is worse than A100 GPU memory, with 3.2x lower per-GPU MTBE for memory errors, (ii) The GPU memory error-recovery mechanisms on H100 GPUs are insufficient to handle the increased memory capacity, (iii) H100 GPUs demonstrate significantly improved GPU hardware resilience over A100 GPUs with respect to critical hardware components, (iv) GPU errors on both A100 and H100 GPUs frequently result in job failures due to the lack of robust recovery mechanisms at the application level, and (v) We project the impact of GPU node availability on larger-scales and find that significant overprovisioning of 5% is necessary to handle GPU failures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f07410f-a34d-499d-a803-46085c44e04dCited by top-tier papers3
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on SuperchipsXinyu Lian, Masahiro Tanaka, Olatunji Ruwase, Minjia ZhangASPLOS 2026 · 6 citations
- GeForge: Hammering GDDR Memory to Forge GPU Page Tables for Fun and ProfitJunpeng Wan, Yanan Guo, Zhi Zhang, Zhuo Li et al.S&P 2026 · 4 citations
- Scaling Out Chip Interconnect Networks with Implicit Sequence NumbersGiyong Jung, Saeid Gorgin, John Kim, Jungrae KimSC 2025 · 2 citations
Builds on8
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang et al.NSDI 2024 · 415 citations
- ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model DevelopmentBorui Wan, Mingji Han, Yiyao Sheng, Yanghua Peng et al.NSDI 2025 · 46 citations
- GPU lifetimes on titan supercomputer: survival analysis and reliabilityGeorge Ostrouchov, Don Maxwell, Rizwan A. Ashraf, Christian Engelmann et al.SC 2020 · 38 citations
- DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language ModelsAvinash Maurya, Robert Underwood, M. Mustafa Rafique, Franck Cappello et al.HPDC 2024 · 34 citations
- GPU-trident: efficient modeling of error propagation in GPU programsAbdul Rehman Anwer, Guanpeng Li, Karthik Pattabiraman, Michael B. Sullivan et al.SC 2020 · 31 citations
Related papers
- TrainMover: An Interruption-Resilient Runtime for ML TrainingChonLam Lao, Jiaqi Gao, Jiamin Cao, Zhipeng Zhang et al.OSDI 2026
- Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated SystemsYonatan Levitt, Richard Barella, Sam Zeltner, Thomas Musta et al.SC 2025 · 1 citation
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev et al.EuroSys 2024 · 23 citations
- EROICA: Online Performance Troubleshooting for Large-scale Model TrainingYu Guan, Zhiyu Yin, Haoyu Chen, Sheng Cheng et al.NSDI 2026 · 1 citation
- G-SEPM: building an accurate and efficient soft error prediction model for GPGPUsHengshan Yue, Xiaohui Wei, Guangli Li, Jianpeng Zhao et al.SC 2021 · 17 citations
