Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNs
John Thorpe, Pengzhan Zhao, Jonathan Eyolfson, Yifan Qiao, Zhihao Jia, Minjia Zhang, Ravi Netravali, Guoqing Harry Xu
摘要
DNN models across many domains continue to grow in size, resulting in high resource requirements for effective training, and unpalatable (and often unaffordable) costs for organizations and research labs across scales. This paper aims to significantly reduce training costs with effective use of preemptible instances, i.e., those that can be obtained at a much cheaper price while idle, but may be preempted whenever requested by priority users. Doing so, however, requires new forms of resiliency and efficiency to cope with the possibility of frequent preemptions -a failure model that is drastically different from the occasional failures in normal cluster settings that existing checkpointing techniques target.
We present Bamboo, a distributed system that tackles these challenges by introducing redundant computations into the training pipeline, i.e., whereby one node performs computations over not only its own layers but also over some layers in its neighbor. Our key insight is that training large models often requires pipeline parallelism where "pipeline bubbles" naturally exist. Bamboo carefully fills redundant computations into these bubbles, providing resilience at a low cost. Across a variety of widely used DNN models, Bamboo outperforms traditional checkpointing by 3.7× in training throughput, and reduces costs by 2.4× compared to a setting where on-demand instances are used.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Can't Be Late: Optimizing Spot Instance Savings under DeadlinesZhanghao Wu, Wei-Lin Chiang, Ziming Mao, Zongheng Yang 等NSDI 2024 · 被引用 40 次
- Oobleck: Resilient Distributed Training of Large Models Using Pipeline TemplatesInsu Jang, Zhenning Yang, Zhen Zhang, Xin Jin 等SOSP 2023 · 被引用 27 次
- Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable ParallelismXinyu Lian, Sam Ade Jacobs, Lev Kurilenko, Masahiro Tanaka 等USENIX ATC 2025 · 被引用 22 次
- SkyServe: Serving AI Models across Regions and Clouds with Spot InstancesZiming Mao, Tian Xia, Zhanghao Wu, Wei-Lin Chiang 等EuroSys 2025 · 被引用 16 次
- ReCycle: Resilient Training of Large DNNs using Pipeline AdaptationSwapnil Gandhi, Mark Zhao, Athinagoras Skiadopoulos, Christos KozyrakisSOSP 2024 · 被引用 13 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee 等OSDI 2020 · 被引用 286 次
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang 等OSDI 2020 · 被引用 260 次
相关 Paper
- Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible InstancesJiangfei Duan, Ziang Song, Xupeng Miao, Xiaoli Xi 等NSDI 2024 · 被引用 54 次
- Machine Learning on Volatile InstancesXiaoxi Zhang, Jianyu Wang, Gauri Joshi, Carlee Joe-WongINFOCOM 2020 · 被引用 17 次
- SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-EfficientMax Ryabinin, Tim Dettmers, Michael Diskin, Alexander BorzunovICML 2023 · 被引用 63 次
- SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUsJin Lee, Zhonghao Chen, Xuhang He, Robert Underwood 等ICML 2026 · 被引用 1 次
- LowDiff: Efficient Frequent Checkpointing via Low-Cost Differential for High-Performance Distributed Training SystemsChenxuan Yao, Feifan Liu, Yuchong Hu, Zhengyu Liu 等SC 2025 · 被引用 3 次
