A nchor : Mitigating GPU Shallow Disruptions with Decoupled Memory
Haoyi Ma, Shiwei Gao, Youmin Chen, Junrong Huang, Youyou Lu, Jiwu Shu
2026年份
摘要
Job failures are frequent in large-scale GPU clusters for LLM workloads, leading to significant resource wastage. The vast majority of these are shallow disruptions (e.g., software errors or updates), where only the worker process crashes while the underlying GPU and OS kernel remain intact. Existing recovery systems, however, are failure-agnostic, relying on heavyweight checkpointing that forces a full, slow state reload even when the data could have survived on the GPU.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev 等EuroSys 2024 · 被引用 23 次
- Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed TrainingXuanyu Wang, Fangcheng Fu, Haoyang Li, Hao Ge 等PPoPP 2026 · 被引用 1 次
- TrainMover: An Interruption-Resilient Runtime for ML TrainingChonLam Lao, Jiaqi Gao, Jiamin Cao, Zhipeng Zhang 等OSDI 2026
- ResiHP: Taming LLM Training Failures with Dynamic Hybrid ParallelismTenghui Ma, Jihu Guo, Wei Gao, Sitian Lu 等HPDC 2026
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang 等NSDI 2024 · 被引用 192 次
