Lune

INFOCOM2026Top-tier venue

Rehabilitating over Recomputing: A Novel Failure Recovery Method for Large Model Training

Zichen Wang, Hongliang Li, Jie Wu, Zhewen Xu, Hairui Zhao, Qi Tian, Haixiao Xu

2026Year

Abstract

Today, large DNN models are often trained in parallel and distributed environments, which inevitably incur high costs for fault tolerance. Existing solutions rely on periodic checkpoints to save training state in run-time and recomputing from the latest checkpoint to recover upon a failure. Despite recent efforts to reduce checkpoint cost, data transfer bandwidth can become a bottleneck, limiting checkpoint frequency, and consequently leading to high recomputing cost. This paper focuses on efficient fault tolerance for large model training and proposes Controlled Predicting-assisted Self-Recovery (CPSR). Instead of resource-consuming recomputation, it features a lightweight predictor, fed by routine checkpoints, to predict training state prior to a failure. The prediction error exhibits as a minor perturbation that can be self-corrected by the training process itself. We propose a quantified model of predicting-based recovery cost during rehabilitation and introduce a novel checkpoint interval problem that seeks to minimize the overall fault tolerance cost. We present a solution to compute the optimal checkpoint interval configuration for a given setting, balancing checkpoint and recovery cost. Extensive testbed experiment data demonstrate that CPSR reduces the recovery cost by 41.66% on average compared with state-of-the-art approaches, while introducing a small GPU memory footprint (less than 200MB).

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines