Lune

SC2025Top-tier venue

Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems

Yonatan Levitt, Richard Barella, Sam Zeltner, Thomas Musta, Lance Cheney, Gustavo Espinosa, Olivier Franza, Balazs Gerofi

2025Year
1Citations

Abstract

As high-performance computing (HPC) systems scale in size, system wide hardware failure rates increase. Historical data from previous large-scale HPC installations illustrate this trend, with the mean time between failures (MTBF) decreasing steadily over the past decade. Recent studies from artificial intelligence and machine-learning (AI/ML) training extrapolate MTBF declining even further for future GPU accelerated systems. As MTBF decreases, mean time to repair (MTTR) becomes more pronounced, highlighting the need for efficient recovery strategies.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get ec58dff3-58be-446f-92c3-897dfb6c903e

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines