Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation Models
Assaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere, Raghuraman Krishnamoorthi, Krishnakumar Nair, Misha Smelyanskiy, Murali Annavaram
Abstract
Checkpoints play an important role in training long running machine learning (ML) models. Checkpoints take a snapshot of an ML model and store it in a non-volatile memory so that they can be used to recover from failures to ensure rapid training progress. In addition, they are used for online training to improve inference prediction accuracy with continuous learning. Given the large and ever-increasing model sizes, checkpoint frequency is often bottlenecked by the storage write bandwidth and capacity. When checkpoints are maintained on remote storage, as is the case with many industrial settings, they are also bottlenecked by network bandwidth. We present Check-N-Run, a scalable checkpointing system for training large ML models at Facebook. While Check-N-Run is applicable to long running ML jobs, we focus on checkpointing recommendation models which are currently the largest ML models with Terabytes of model size. Check-N-Run uses two primary techniques to address the size and bandwidth challenges. First, it applies differential checkpointing, which tracks and checkpoints the modified part of the model. Differential checkpointing is particularly valuable in the context of recommendation models where only a fraction of the model (stored as embedding tables) is updated on each iteration. Second, Check-N-Run leverages quantization techniques to significantly reduce the checkpoint size, without degrading training accuracy. These techniques allow Check-N-Run to reduce the required write bandwidth by 6-17× and the required capacity by 2.5-8× on real-world models at Facebook, and thereby significantly improve checkpoint capabilities while reducing the total cost of ownership.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c9256bb-3dc9-4838-8e41-d6f32a877b0bCited by top-tier papers35
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang et al.NSDI 2024 · 192 citations
- GEMINI: Fast Failure Recovery in Distributed Training with In-Memory CheckpointsZhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang et al.SOSP 2023 · 61 citations
- DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM ServingFoteini Strati, Sara McAllister, Amar Phanishayee, Jakub Tarnawski et al.ICML 2024 · 59 citations
- ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model DevelopmentBorui Wan, Mingji Han, Yiyao Sheng, Yanghua Peng et al.NSDI 2025 · 46 citations
- TapFinger: Task Placement and Fine-Grained Resource Allocation for Edge Machine LearningYihong Li, Tianyu Zeng, Xiaoxi Zhang, Jingpu Duan et al.INFOCOM 2023 · 39 citations
Builds on1
Related papers
- LowDiff: Efficient Frequent Checkpointing via Low-Cost Differential for High-Performance Distributed Training SystemsChenxuan Yao, Feifan Liu, Yuchong Hu, Zhengyu Liu et al.SC 2025 · 3 citations
- QuickUpdate: a Real-Time Personalization System for Large-Scale Recommendation ModelsKiran Kumar Matam, Hani Ramezani, Fan Wang, Zeliang Chen et al.NSDI 2024 · 13 citations
- IncrCP: Decomposing and Orchestrating Incremental Checkpoints for Effective Recommendation Model TrainingQingyin Lin, Jiangsu Du, Rui Li, Zhiguang Chen et al.VLDB 2025 · 2 citations
- Efficient Fault Tolerance for Recommendation Model Training via Erasure CodingTianyu Zhang, Kaige Liu, Jack Kosaian, Juncheng Yang et al.VLDB 2023 · 10 citations
- PCcheck: Persistent Concurrent Checkpointing for MLFoteini Strati, Michal Friedman, Ana KlimovicASPLOS 2025 · 11 citations
