Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
Shengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan, Ziheng Chen, Phuong Cao, Gregory H. Bauer, Brett M. Bode, Catello Di Martino, Saurabh Jha, Chandra Narayanaswami, Daby Sow
摘要
This study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years of operational data (11.7 million GPU hours) on GPU errors. Our major findings include: (i) H100 GPU memory resilience is worse than A100 GPU memory, with 3.2x lower per-GPU MTBE for memory errors, (ii) The GPU memory error-recovery mechanisms on H100 GPUs are insufficient to handle the increased memory capacity, (iii) H100 GPUs demonstrate significantly improved GPU hardware resilience over A100 GPUs with respect to critical hardware components, (iv) GPU errors on both A100 and H100 GPUs frequently result in job failures due to the lack of robust recovery mechanisms at the application level, and (v) We project the impact of GPU node availability on larger-scales and find that significant overprovisioning of 5% is necessary to handle GPU failures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on SuperchipsXinyu Lian, Masahiro Tanaka, Olatunji Ruwase, Minjia ZhangASPLOS 2026 · 被引用 6 次
- GeForge: Hammering GDDR Memory to Forge GPU Page Tables for Fun and ProfitJunpeng Wan, Yanan Guo, Zhi Zhang, Zhuo Li 等S&P 2026 · 被引用 4 次
- Scaling Out Chip Interconnect Networks with Implicit Sequence NumbersGiyong Jung, Saeid Gorgin, John Kim, Jungrae KimSC 2025 · 被引用 2 次
它引用的顶会 Paper8
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 等NSDI 2024 · 被引用 415 次
- ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model DevelopmentBorui Wan, Mingji Han, Yiyao Sheng, Yanghua Peng 等NSDI 2025 · 被引用 46 次
- GPU lifetimes on titan supercomputer: survival analysis and reliabilityGeorge Ostrouchov, Don Maxwell, Rizwan A. Ashraf, Christian Engelmann 等SC 2020 · 被引用 38 次
- DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language ModelsAvinash Maurya, Robert Underwood, M. Mustafa Rafique, Franck Cappello 等HPDC 2024 · 被引用 34 次
- GPU-trident: efficient modeling of error propagation in GPU programsAbdul Rehman Anwer, Guanpeng Li, Karthik Pattabiraman, Michael B. Sullivan 等SC 2020 · 被引用 31 次
相关 Paper
- TrainMover: An Interruption-Resilient Runtime for ML TrainingChonLam Lao, Jiaqi Gao, Jiamin Cao, Zhipeng Zhang 等OSDI 2026
- Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated SystemsYonatan Levitt, Richard Barella, Sam Zeltner, Thomas Musta 等SC 2025 · 被引用 1 次
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev 等EuroSys 2024 · 被引用 23 次
- EROICA: Online Performance Troubleshooting for Large-scale Model TrainingYu Guan, Zhiyu Yin, Haoyu Chen, Sheng Cheng 等NSDI 2026 · 被引用 1 次
- G-SEPM: building an accurate and efficient soft error prediction model for GPGPUsHengshan Yue, Xiaohui Wei, Guangli Li, Jianpeng Zhao 等SC 2021 · 被引用 17 次
