A Self-updating Checkpointing and Fast Failure Recovery System for Distributed LLM Training
Leyi Ye, Zhiyi Yao, Boliang Liu, Yuedong Xu, Liang Du, Zeng Chuxuan, Ling Deng, Jessie Hui Wang
摘要
Checkpointing is the most prominent mechanism for large-scale deep learning training to ensure reliability, against frequent failures. Existing checkpointing systems suffer from significant communication overhead of transmitting checkpoints to remote storage or remote CPU memory. This leads to checkpointing-induced stalls, which then slows down the training. This paper presents Zoetic, a novel checkpointing system that performs checkpointing at every iteration without impairing training efficiency. The key idea behind Zoetic is to enable a machine to update remote checkpoints in its CPU memory by leveraging regular training traffic, inherent in gradient synchronization of data-parallelism. In contrast to traditional checkpoint transmission, Zoetic reduces checkpointing traffic to one-sixth or even zero, effectively eliminating checkpointing-induced stalls. We orchestrate the transmission of gradients from GPU HBM to CPU memory to avoid interfering with regular training traffic, and design an enhanced CPU checkpoint optimizer to overlap its operation with GPU training computations. We further propose a fast recovery method, which utilizes direct remote GPU access and non-serialized transmission to reduce recovery latency. Experimental results show that Zoetic achieves iteration-wise checkpointing without prolonging training time, while existing approaches slow down training by up to 41.5%. Additionally, Zoetic exhibits an 11.7 × faster checkpoint loading efficiency.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- GEMINI: Fast Failure Recovery in Distributed Training with In-Memory CheckpointsZhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang 等SOSP 2023 · 被引用 61 次
- FlowCheck: Decoupling Checkpointing and Training of Large-Scale ModelsZimeng Huang, Hao Nie, Haonan Jia, Bo Jiang 等EuroSys 2025 · 被引用 3 次
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev 等EuroSys 2024 · 被引用 23 次
- LowDiff: Efficient Frequent Checkpointing via Low-Cost Differential for High-Performance Distributed Training SystemsChenxuan Yao, Feifan Liu, Yuchong Hu, Zhengyu Liu 等SC 2025 · 被引用 3 次
- Checkmate: Zero Performance Overhead Model Checkpointing via Network Gradient ReplicationAnkit Bhardwaj, Weiyang Wang, Jeremy Carin, Adam Belay 等NSDI 2026 · 被引用 3 次
