Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed Training
Xuanyu Wang, Fangcheng Fu, Haoyang Li, Hao Ge, Sheng Lin, Jiawen Niu, Bin Cui
2026Year
1Citations
1Top-tier citations
Abstract
Distributed deep learning (DL) training faces instability from GPU/node failures of multi-GPU clusters, necessitating robust fault recovery from model checkpoints. However, we find that existing works only considers node failures but fails to handle partial GPU unavailability, and suffers from inefficient model checkpointing saving and loading, particularly when the GPU availability changes.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fffdf742-8773-45da-b88e-2cf2174dbb21Cited by top-tier papers1
Ask how each one uses itRelated papers
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev et al.EuroSys 2024 · 23 citations
- A nchor : Mitigating GPU Shallow Disruptions with Decoupled MemoryHaoyi Ma, Shiwei Gao, Youmin Chen, Junrong Huang et al.SOSP 2026
- LowDiff: Efficient Frequent Checkpointing via Low-Cost Differential for High-Performance Distributed Training SystemsChenxuan Yao, Feifan Liu, Yuchong Hu, Zhengyu Liu et al.SC 2025 · 3 citations
- Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor CollectionsMarcel Wagenländer, Guo Li, Bo Zhao, Luo Mai et al.SOSP 2024 · 8 citations
- CheckFreq: Frequent, Fine-Grained DNN CheckpointingJayashree Mohan, Amar Phanishayee, Vijay ChidambaramFAST 2021 · 175 citations
