Lune

INFOCOM2026顶会

A Self-updating Checkpointing and Fast Failure Recovery System for Distributed LLM Training

Leyi Ye, Zhiyi Yao, Boliang Liu, Yuedong Xu, Liang Du, Zeng Chuxuan, Ling Deng, Jessie Hui Wang

2026年份

摘要

Checkpointing is the most prominent mechanism for large-scale deep learning training to ensure reliability, against frequent failures. Existing checkpointing systems suffer from significant communication overhead of transmitting checkpoints to remote storage or remote CPU memory. This leads to checkpointing-induced stalls, which then slows down the training. This paper presents Zoetic, a novel checkpointing system that performs checkpointing at every iteration without impairing training efficiency. The key idea behind Zoetic is to enable a machine to update remote checkpoints in its CPU memory by leveraging regular training traffic, inherent in gradient synchronization of data-parallelism. In contrast to traditional checkpoint transmission, Zoetic reduces checkpointing traffic to one-sixth or even zero, effectively eliminating checkpointing-induced stalls. We orchestrate the transmission of gradients from GPU HBM to CPU memory to avoid interfering with regular training traffic, and design an enhanced CPU checkpoint optimizer to overlap its operation with GPU training computations. We further propose a fast recovery method, which utilizes direct remote GPU access and non-serialized transmission to reduce recovery latency. Experimental results show that Zoetic achieves iteration-wise checkpointing without prolonging training time, while existing approaches slow down training by up to 41.5%. Additionally, Zoetic exhibits an 11.7 × faster checkpoint loading efficiency.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get f0d570e5-55a2-4cb6-9a44-2c5efaf79a1a

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖