A nchor : Mitigating GPU Shallow Disruptions with Decoupled Memory
Haoyi Ma, Shiwei Gao, Youmin Chen, Junrong Huang, Youyou Lu, Jiwu Shu
2026Year
Abstract
Job failures are frequent in large-scale GPU clusters for LLM workloads, leading to significant resource wastage. The vast majority of these are shallow disruptions (e.g., software errors or updates), where only the worker process crashes while the underlying GPU and OS kernel remain intact. Existing recovery systems, however, are failure-agnostic, relying on heavyweight checkpointing that forces a full, slow state reload even when the data could have survived on the GPU.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e3725449-6e8c-4842-b4e1-c1bbc58bc16dRelated papers
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev et al.EuroSys 2024 · 23 citations
- Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed TrainingXuanyu Wang, Fangcheng Fu, Haoyang Li, Hao Ge et al.PPoPP 2026 · 1 citation
- TrainMover: An Interruption-Resilient Runtime for ML TrainingChonLam Lao, Jiaqi Gao, Jiamin Cao, Zhipeng Zhang et al.OSDI 2026
- ResiHP: Taming LLM Training Failures with Dynamic Hybrid ParallelismTenghui Ma, Jihu Guo, Wei Gao, Sitian Lu et al.HPDC 2026
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang et al.NSDI 2024 · 192 citations
