GPU-Enabled Asynchronous Multi-level Checkpoint Caching and Prefetching
Avinash Maurya, M. Mustafa Rafique, Thierry Tonellot, Hussain J. AlSalem, Franck Cappello, Bogdan Nicolae
摘要
Checkpointing is an I/O intensive operation increasingly used by High-Performance Computing (HPC) applications to revisit previous intermediate datasets at scale. Unlike the case of resilience, where only the last checkpoint is needed for application restart and rarely accessed to recover from failures, in this scenario, it is important to optimize frequent reads and writes of an entire history of checkpoints. State-of-the-art checkpointing approaches often rely on asynchronous multi-level techniques to hide I/O overheads by writing to fast local tiers (e.g. an SSD) and asynchronously flushing to slower, potentially remote tiers (e.g. a parallel file system) in the background, while the application keeps running. However, such approaches have two limitations. First, despite the fact that HPC infrastructures routinely rely on accelerators (e.g. GPUs), and therefore a majority of the checkpoints involve GPU memory, efficient asynchronous data movement between the GPU memory and host memory is lagging behind. Second, revisiting previous data often involves predictable access patterns, which are not exploited to accelerate read operations. In this paper, we address these limitations by proposing a scalable and asynchronous multi-level checkpointing approach optimized for both reading and writing of an arbitrarily long history of checkpoints. Our approach exploits GPU memory as a first-class citizen in the multi-level storage hierarchy to enable informed caching and prefetching of checkpoints by leveraging foreknowledge about the access order passed by the application as hints. Our evaluation using a variety of scenarios under I/O concurrency shows up to 74× faster checkpoint and restore throughput as compared to the state-of-art runtime and optimized unified virtual memory (UVM) based prefetching strategies and at least 2× shorter I/O wait time for the application across various workloads and configurations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language ModelsAvinash Maurya, Robert Underwood, M. Mustafa Rafique, Franck Cappello 等HPDC 2024 · 被引用 34 次
- IncrCP: Decomposing and Orchestrating Incremental Checkpoints for Effective Recommendation Model TrainingQingyin Lin, Jiangsu Du, Rui Li, Zhiguang Chen 等VLDB 2025 · 被引用 2 次
- SIVF: GPU-Resident IVF Index for Streaming Vector AnalyticsDongfang ZhaoHPDC 2026
它引用的顶会 Paper3
- CheckFreq: Frequent, Fine-Grained DNN CheckpointingJayashree Mohan, Amar Phanishayee, Vijay ChidambaramFAST 2021 · 被引用 175 次
- Clairvoyant prefetching for distributed machine learning I/ONikoli Dryden, Roman Böhringer, Tal Ben-Nun, Torsten HoeflerSC 2021 · 被引用 60 次
- Canary: Fault-Tolerant FaaS for Stateful Time-Sensitive ApplicationsMoiz Arif, Kevin Assogba, M. Mustafa RafiqueSC 2022 · 被引用 9 次
相关 Paper
- GPU Checkpoint/Restore Made Fast and LightweightShaoxun Zeng, Tingxu Ren, Jiwu Shu, Youyou LuFAST 2026 · 被引用 5 次
- CoPilotIO: CPU as a Co-Pilot for GPU I/O to Free GPU ComputeGuanyi Chen, Qi Chen, Shu Yin, Jian ZhangOSDI 2026
- PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated SpeculationXingda Wei, Zhuobin Huang, Tianle Sun, Yingyi Hao 等SOSP 2025 · 被引用 3 次
- Asynchrony and GPUs: Bridging this Dichotomy for I/O with AGIOJihoon Han, Anand Sivasubramaniam, Chia-Hao Chang, Vikram Sharma Mailthody 等ASPLOS 2026 · 被引用 1 次
- GMT: GPU Orchestrated Memory Tiering for the Big Data EraChia-Hao Chang, Jihoon Han, Anand Sivasubramaniam, Vikram Sharma Mailthody 等ASPLOS 2024 · 被引用 11 次
