A Self-updating Checkpointing and Fast Failure Recovery System for Distributed LLM Training
Leyi Ye, Zhiyi Yao, Boliang Liu, Yuedong Xu, Liang Du, Zeng Chuxuan, Ling Deng, Jessie Hui Wang
Abstract
Checkpointing is the most prominent mechanism for large-scale deep learning training to ensure reliability, against frequent failures. Existing checkpointing systems suffer from significant communication overhead of transmitting checkpoints to remote storage or remote CPU memory. This leads to checkpointing-induced stalls, which then slows down the training. This paper presents Zoetic, a novel checkpointing system that performs checkpointing at every iteration without impairing training efficiency. The key idea behind Zoetic is to enable a machine to update remote checkpoints in its CPU memory by leveraging regular training traffic, inherent in gradient synchronization of data-parallelism. In contrast to traditional checkpoint transmission, Zoetic reduces checkpointing traffic to one-sixth or even zero, effectively eliminating checkpointing-induced stalls. We orchestrate the transmission of gradients from GPU HBM to CPU memory to avoid interfering with regular training traffic, and design an enhanced CPU checkpoint optimizer to overlap its operation with GPU training computations. We further propose a fast recovery method, which utilizes direct remote GPU access and non-serialized transmission to reduce recovery latency. Experimental results show that Zoetic achieves iteration-wise checkpointing without prolonging training time, while existing approaches slow down training by up to 41.5%. Additionally, Zoetic exhibits an 11.7 × faster checkpoint loading efficiency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f0d570e5-55a2-4cb6-9a44-2c5efaf79a1aRelated papers
- GEMINI: Fast Failure Recovery in Distributed Training with In-Memory CheckpointsZhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang et al.SOSP 2023 · 61 citations
- FlowCheck: Decoupling Checkpointing and Training of Large-Scale ModelsZimeng Huang, Hao Nie, Haonan Jia, Bo Jiang et al.EuroSys 2025 · 3 citations
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev et al.EuroSys 2024 · 23 citations
- LowDiff: Efficient Frequent Checkpointing via Low-Cost Differential for High-Performance Distributed Training SystemsChenxuan Yao, Feifan Liu, Yuchong Hu, Zhengyu Liu et al.SC 2025 · 3 citations
- Checkmate: Zero Performance Overhead Model Checkpointing via Network Gradient ReplicationAnkit Bhardwaj, Weiyang Wang, Jeremy Carin, Adam Belay et al.NSDI 2026 · 3 citations
