PCcheck: Persistent Concurrent Checkpointing for ML
Foteini Strati, Michal Friedman, Ana Klimovic
Abstract
Training large-scale machine learning (ML) models is expensive and time-intensive, consuming many hardware accelerators for days or weeks. As the scale of hardware deployments and training time continue to grow, the probability of failures also increases. The desire to use cheaper cloud resources, such as spot VMs, to lower costs also dramatically increases the frequency of failures. The standard approach to deal with failures is to periodically pause training and checkpoint model parameters to persistent storage. Unfortunately, today's checkpointing mechanisms introduce high overhead when applied at high frequencies, yet frequent checkpointing is necessary to avoid long recovery times.
We present a concurrent checkpointing mechanism, PCcheck, that allows frequent checkpointing with minimal overhead. Our framework supports persisting checkpoints to SSD and persistent main memory (PMEM) for both singlemachine and distributed settings. PCcheck enables checkpointing as frequently as every 10 iterations for detailed monitoring and fast recovery times in case of failures, while maintaining minimal (3%) overhead on training throughput.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a1da35d-9c19-4929-892b-539ac0b65120Cited by top-tier papers4
- AdaCheck: An Adaptive Checkpointing System for Efficient LLM Training with Redundancy UtilizationWeijie Liu, Shengwei Li, Zhiquan Lai, Keshi Ge et al.FAST 2026 · 3 citations
- Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed ClustersFoteini Strati, Zhendong Zhang, George Manos, Ixeia Sánchez Périz et al.SOSP 2025 · 2 citations
- Sparse Checkpointing for Fast and Reliable MoE TrainingSwapnil Gandhi, Christos KozyrakisNSDI 2026
- ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless CompressionRuibo Fan, Xiangrui Yu, Xinglin Pan, Zeyu Li et al.ASPLOS 2026
Builds on16
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- An Empirical Guide to the Behavior and Use of Scalable Persistent MemoryJian Yang, Juno Kim, Morteza Hoseinzadeh, Joseph Izraelevitz et al.FAST 2020 · 470 citations
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang et al.NSDI 2024 · 415 citations
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger et al.OSDI 2021 · 258 citations
- CheckFreq: Frequent, Fine-Grained DNN CheckpointingJayashree Mohan, Amar Phanishayee, Vijay ChidambaramFAST 2021 · 175 citations
Related papers
- DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language ModelsAvinash Maurya, Robert Underwood, M. Mustafa Rafique, Franck Cappello et al.HPDC 2024 · 34 citations
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev et al.EuroSys 2024 · 23 citations
- GEMINI: Fast Failure Recovery in Distributed Training with In-Memory CheckpointsZhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang et al.SOSP 2023 · 61 citations
- LowDiff: Efficient Frequent Checkpointing via Low-Cost Differential for High-Performance Distributed Training SystemsChenxuan Yao, Feifan Liu, Yuchong Hu, Zhengyu Liu et al.SC 2025 · 3 citations
- Rehabilitating over Recomputing: A Novel Failure Recovery Method for Large Model TrainingZichen Wang, Hongliang Li, Jie Wu, Zhewen Xu et al.INFOCOM 2026
