Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNs
John Thorpe, Pengzhan Zhao, Jonathan Eyolfson, Yifan Qiao, Zhihao Jia, Minjia Zhang, Ravi Netravali, Guoqing Harry Xu
Abstract
DNN models across many domains continue to grow in size, resulting in high resource requirements for effective training, and unpalatable (and often unaffordable) costs for organizations and research labs across scales. This paper aims to significantly reduce training costs with effective use of preemptible instances, i.e., those that can be obtained at a much cheaper price while idle, but may be preempted whenever requested by priority users. Doing so, however, requires new forms of resiliency and efficiency to cope with the possibility of frequent preemptions -a failure model that is drastically different from the occasional failures in normal cluster settings that existing checkpointing techniques target.
We present Bamboo, a distributed system that tackles these challenges by introducing redundant computations into the training pipeline, i.e., whereby one node performs computations over not only its own layers but also over some layers in its neighbor. Our key insight is that training large models often requires pipeline parallelism where "pipeline bubbles" naturally exist. Bamboo carefully fills redundant computations into these bubbles, providing resilience at a low cost. Across a variety of widely used DNN models, Bamboo outperforms traditional checkpointing by 3.7× in training throughput, and reduces costs by 2.4× compared to a setting where on-demand instances are used.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57139a26-db9d-4df0-a3d1-e3a80cf29302Cited by top-tier papers20
- Can't Be Late: Optimizing Spot Instance Savings under DeadlinesZhanghao Wu, Wei-Lin Chiang, Ziming Mao, Zongheng Yang et al.NSDI 2024 · 40 citations
- Oobleck: Resilient Distributed Training of Large Models Using Pipeline TemplatesInsu Jang, Zhenning Yang, Zhen Zhang, Xin Jin et al.SOSP 2023 · 27 citations
- Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable ParallelismXinyu Lian, Sam Ade Jacobs, Lev Kurilenko, Masahiro Tanaka et al.USENIX ATC 2025 · 22 citations
- SkyServe: Serving AI Models across Regions and Clouds with Spot InstancesZiming Mao, Tian Xia, Zhanghao Wu, Wei-Lin Chiang et al.EuroSys 2025 · 16 citations
- ReCycle: Resilient Training of Large DNNs using Pipeline AdaptationSwapnil Gandhi, Mark Zhao, Athinagoras Skiadopoulos, Christos KozyrakisSOSP 2024 · 13 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee et al.OSDI 2020 · 286 citations
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
Related papers
- Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible InstancesJiangfei Duan, Ziang Song, Xupeng Miao, Xiaoli Xi et al.NSDI 2024 · 54 citations
- Machine Learning on Volatile InstancesXiaoxi Zhang, Jianyu Wang, Gauri Joshi, Carlee Joe-WongINFOCOM 2020 · 17 citations
- SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-EfficientMax Ryabinin, Tim Dettmers, Michael Diskin, Alexander BorzunovICML 2023 · 63 citations
- SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUsJin Lee, Zhonghao Chen, Xuhang He, Robert Underwood et al.ICML 2026 · 1 citation
- LowDiff: Efficient Frequent Checkpointing via Low-Cost Differential for High-Performance Distributed Training SystemsChenxuan Yao, Feifan Liu, Yuchong Hu, Zhengyu Liu et al.SC 2025 · 3 citations
