GPU Checkpoint/Restore Made Fast and Lightweight
Shaoxun Zeng, Tingxu Ren, Jiwu Shu, Youyou Lu
Abstract
System-level GPU checkpoint/restore (C/R) enables several critical features such as elastic scaling, task switching, and fault tolerance, for modern GPU workloads in a unified and application-transparent manner. However, existing approaches present fundamental limitations: they fail to simultaneously achieve low C/R latency and low overhead imposed on normal GPU execution, while also lacking efficient support for incremental checkpointing. We propose GCR, a GPU checkpoint/restore system that addresses all these limitations simultaneously. GCR employs a hybrid C/R scheme through control/data separation to deliver low C/R latency and negligible overhead imposed on normal GPU execution. To efficiently support incremental checkpointing, GCR introduces shadow execution on the CPU to reduce the overhead of dirty buffer identification, utilizing dirty templates for both lightweight CPU shadow execution and identification at a fine-grained instruction level.
Our evaluations demonstrate that GCR reduces GPU checkpointing latency by 72.1% and 63.6% compared to cuda-ckpt (NVIDIA's official solution) and PhOS (the current state-ofthe-art), respectively, and restoration latency by 54.2% and 87.1%, while imposing negligible overhead (less than 1%). GCR also supports efficient incremental checkpointing, which reduces checkpoint sizes by 86.6% and latency by 43.8%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7cb1f00-ea65-4c86-aef8-e503980c0e13Builds on13
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah et al.ISCA 2024 · 282 citations
- Balancing efficiency and fairness in heterogeneous GPU clusters for deep learningShubham Chaudhary, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra et al.EuroSys 2020 · 135 citations
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete et al.OSDI 2024 · 125 citations
Related papers
- PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated SpeculationXingda Wei, Zhuobin Huang, Tianle Sun, Yingyi Hao et al.SOSP 2025 · 3 citations
- CRAC: checkpoint-restart architecture for CUDA with streams and UVMTwinkle Jain, Gene CoopermanSC 2020 · 16 citations
- GPU-Enabled Asynchronous Multi-level Checkpoint Caching and PrefetchingAvinash Maurya, M. Mustafa Rafique, Thierry Tonellot, Hussain J. AlSalem et al.HPDC 2023 · 13 citations
- CheckFreq: Frequent, Fine-Grained DNN CheckpointingJayashree Mohan, Amar Phanishayee, Vijay ChidambaramFAST 2021 · 175 citations
- libcrpm: improving the checkpoint performance of NVMFeng Ren, Kang Chen, Yongwei WuDAC 2022 · 2 citations
