SC2020Top-tier venue
CRAC: checkpoint-restart architecture for CUDA with streams and UVM
Twinkle Jain, Gene Cooperman
Abstract
The share of the top 500 supercomputers with NVIDIA GPUs is now over 25% and continues to grow. While fault tolerance is a critical issue for supercomputing, there does not currently exist an efficient, scalable solution for CUDA applications on NVIDIA GPUs. CRAC (Checkpoint-Restart Architecture for CUDA) is a new checkpoint-restart solution for fault tolerance that supports the full range of CUDA applications. CRAC combines: low runtime overhead (approximately 1% or less); fast checkpoint-restart; support for scalable CUDA streams (for efficient usage of all of the thousands of GPU cores); and support for the full features of Unified Virtual Memory (eliminating the programmer's burden of migrating memory between device and host). CRAC achieves its flexible architecture by segregating application code (checkpointed) and its external GPU communication via non-reentrant CUDA libraries (not checkpointed) within a single process's memory. This eliminates the high overhead of inter-process communication in earlier approaches, and has fewer limitations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e77e3125-0abc-4486-9c03-728cb75c99e4Cited by top-tier papers3
- ElasticNotebook: Enabling Live Migration for Computational NotebooksZhaoheng Li, Pranav Gor, Rahul Prabhu, Hui Yu et al.VLDB 2024 · 12 citations
- Kishu: Time-Traveling for Computational NotebooksZhaoheng Li, Supawit Chockchowwat, Areet Sheth, Yongjoo Park et al.VLDB 2025 · 11 citations
- GPU Checkpoint/Restore Made Fast and LightweightShaoxun Zeng, Tingxu Ren, Jiwu Shu, Youyou LuFAST 2026 · 5 citations
Related papers
- SuperCollider: Scalable and Effective Data Race Detection for CUDAMark Stephenson, Sana Damani, Mohamed Tarek Ibn Ziad, Anis Ladram et al.PLDI 2026
- AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency AnalysisXiang Fu, Weiping Zhang, Shiman Meng, Xin Huang et al.SC 2024 · 1 citation
- PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated SpeculationXingda Wei, Zhuobin Huang, Tianle Sun, Yingyi Hao et al.SOSP 2025 · 3 citations
- Enabling Software Resilience in GPGPU Applications via Partial Thread ProtectionLishan Yang, Bin Nie, Adwait Jog, Evgenia SmirniICSE 2021 · 25 citations
- CheckFreq: Frequent, Fine-Grained DNN CheckpointingJayashree Mohan, Amar Phanishayee, Vijay ChidambaramFAST 2021 · 175 citations
