Lune

INFOCOM2026顶会

Rehabilitating over Recomputing: A Novel Failure Recovery Method for Large Model Training

Zichen Wang, Hongliang Li, Jie Wu, Zhewen Xu, Hairui Zhao, Qi Tian, Haixiao Xu

2026年份

摘要

Today, large DNN models are often trained in parallel and distributed environments, which inevitably incur high costs for fault tolerance. Existing solutions rely on periodic checkpoints to save training state in run-time and recomputing from the latest checkpoint to recover upon a failure. Despite recent efforts to reduce checkpoint cost, data transfer bandwidth can become a bottleneck, limiting checkpoint frequency, and consequently leading to high recomputing cost. This paper focuses on efficient fault tolerance for large model training and proposes Controlled Predicting-assisted Self-Recovery (CPSR). Instead of resource-consuming recomputation, it features a lightweight predictor, fed by routine checkpoints, to predict training state prior to a failure. The prediction error exhibits as a minor perturbation that can be self-corrected by the training process itself. We propose a quantified model of predicting-based recovery cost during rehabilitation and introduce a novel checkpoint interval problem that seeks to minimize the overall fault tolerance cost. We present a solution to compute the optimal checkpoint interval configuration for a given setting, balancing checkpoint and recovery cost. Extensive testbed experiment data demonstrate that CPSR reduces the recovery cost by 41.66% on average compared with state-of-the-art approaches, while introducing a small GPU memory footprint (less than 200MB).

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖