FlowCheck: Decoupling Checkpointing and Training of Large-Scale Models
Zimeng Huang, Hao Nie, Haonan Jia, Bo Jiang, Junchen Guo, Jianyuan Lu, Rong Wen, Biao Lyu, Shunmin Zhu, Xinbing Wang
Abstract
Checkpointing is becoming a hotspot of interest in both academia and industry as the primary fault-tolerance method for large model training. However, existing checkpoint designs are tightly coupled with the training process, leading to interruptions that reduce overall training efficiency. To reduce the impact of checkpoints on training, this paper presents FlowCheck, a novel checkpointing system that decouples checkpoint operations from the training process, enabling checkpoint saving without blocking the training. Specifically, FlowCheck updates the checkpoints by extracting complete gradient information from the network traffic of normal training. FlowCheck deploys a traffic-mirroring network to support this design. To utilize mirrored traffic for checkpointing operations, two key challenges need to be addressed. First, we need to achieve precise identification and extraction of gradient packets from training traffic. Second, the transmission on the mirror link is unreliable due to its inability to trigger retransmission upon packet loss. Through two key designs: (1) packet-counting-based traffic identification, and (2) packet redundancy recovery mechanism, FlowCheck implements an efficient checkpointing system using the existing training network and solves the above two challenges. Experiments and estimations verify that FlowCheck achieves checkpoint operations with zero impact on training, and demonstrate that FlowCheck achieves over 98% effective training time under practical fault conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4db6e1d2-bc34-48ce-b0b4-02f18aabbcf6Cited by top-tier papers1
Ask how each one uses itBuilds on9
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- DAPPLE: a pipelined data parallel approach for training large modelsShiqing Fan, Yi Rong, Chen Meng, Zongyan Cao et al.PPoPP 2021 · 224 citations
- CheckFreq: Frequent, Fine-Grained DNN CheckpointingJayashree Mohan, Amar Phanishayee, Vijay ChidambaramFAST 2021 · 175 citations
- GEMINI: Fast Failure Recovery in Distributed Training with In-Memory CheckpointsZhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang et al.SOSP 2023 · 61 citations
Related papers
- A Self-updating Checkpointing and Fast Failure Recovery System for Distributed LLM TrainingLeyi Ye, Zhiyi Yao, Boliang Liu, Yuedong Xu et al.INFOCOM 2026
- Checkmate: Zero Performance Overhead Model Checkpointing via Network Gradient ReplicationAnkit Bhardwaj, Weiyang Wang, Jeremy Carin, Adam Belay et al.NSDI 2026 · 3 citations
- Rehabilitating over Recomputing: A Novel Failure Recovery Method for Large Model TrainingZichen Wang, Hongliang Li, Jie Wu, Zhewen Xu et al.INFOCOM 2026
- ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model DevelopmentBorui Wan, Mingji Han, Yiyao Sheng, Yanghua Peng et al.NSDI 2025 · 46 citations
- Sparse Checkpointing for Fast and Reliable MoE TrainingSwapnil Gandhi, Christos KozyrakisNSDI 2026
