IncrCP: Decomposing and Orchestrating Incremental Checkpoints for Effective Recommendation Model Training
Qingyin Lin, Jiangsu Du, Rui Li, Zhiguang Chen, Wenguang Chen, Nong Xiao
摘要
Training large models for modern recommendation systems requires a substantial number of computational devices and extended periods. Since it is essential to store model checkpoints throughout the training progress for accuracy debugging or mitigating potential failures, checkpointing systems are widely used. However, given that recommendation models can scale to hundreds of gigabytes or more, existing solutions often introduce significant overhead in terms of both storage and I/O. In this paper, we present IncrCP, a checkpointing system specifically designed for recommendation models. Given that only a small fraction of model parameters are modified in each iteration, IncrCP creatively leverages the incremental checkpointing strategy and overcomes the inherent slow recovery problem. To support recovering all states throughout the training process, while also ensuring efficient storage utilization and rapid recovery, IncrCP proposes the 2-D chunk approach. It proactively records changed parameters in the training process as well as their indexes, extracts parameters according to duplicated indexes as independent chunk files, and then orchestrates these chunks in the 2-dimensional linked list. In this way, IncrCP achieves fast recovery by loading less unnecessary parameters and performing less deduplication during recovery. Furthermore, IncrCP includes a selective extraction approach to reduce I/O by avoiding worthless extractions and a concatenate approach to reduce random disk access when recovery. Evaluations show that IncrCP achieves up to 6.6× recovery speedup compared to the naive incremental strategy and saves storage space by 60.4% with slight overhead compared to another recovery-friendly strategy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- CheckFreq: Frequent, Fine-Grained DNN CheckpointingJayashree Mohan, Amar Phanishayee, Vijay ChidambaramFAST 2021 · 被引用 175 次
- GEMINI: Fast Failure Recovery in Distributed Training with In-Memory CheckpointsZhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang 等SOSP 2023 · 被引用 61 次
- Bagpipe: Accelerating Deep Recommendation Model TrainingSaurabh Agarwal, Chengpo Yan, Ziyi Zhang, Shivaram VenkataramanSOSP 2023 · 被引用 17 次
- EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding TableZheng Wang, Yuke Wang, Boyuan Feng, Dheevatsa Mudigere 等SC 2022 · 被引用 15 次
- GPU-Enabled Asynchronous Multi-level Checkpoint Caching and PrefetchingAvinash Maurya, M. Mustafa Rafique, Thierry Tonellot, Hussain J. AlSalem 等HPDC 2023 · 被引用 13 次
相关 Paper
- LowDiff: Efficient Frequent Checkpointing via Low-Cost Differential for High-Performance Distributed Training SystemsChenxuan Yao, Feifan Liu, Yuchong Hu, Zhengyu Liu 等SC 2025 · 被引用 3 次
- Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation ModelsAssaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere 等NSDI 2022
- Efficient Fault Tolerance for Recommendation Model Training via Erasure CodingTianyu Zhang, Kaige Liu, Jack Kosaian, Juncheng Yang 等VLDB 2023 · 被引用 10 次
- PCcheck: Persistent Concurrent Checkpointing for MLFoteini Strati, Michal Friedman, Ana KlimovicASPLOS 2025 · 被引用 11 次
- Kraken: memory-efficient continual learning for large-scale real-time recommendationsMinhui Xie, Kai Ren, Youyou Lu, Guangxu Yang 等SC 2020 · 被引用 38 次
