Bonsai: Efficient and Optimal Automatic Tensor Rematerialization for Memory-Constrained DNN Training
Dat Nguyen, Vasudha Devarakonda, Anxiao Jiang, Khanh Nguyen
摘要
GPU memory is increasingly the primary bottleneck in scaling deep neural network (DNN) training, where the activation tensors footprint of a model may exceed the memory capacity. Tensor recomputation is a powerful technique that trades additional computation for reduced peak memory usage. However, existing approaches face a fundamental tension between performance optimality and computational scalability. On the one hand, solvers leverage Integer Linear Programming (ILP) to provide mathematically optimal solutions but suffer from the combinatorial explosion of the search space and thus become intractable for modern DNN models. On the other hand, heuristics-based approaches achieve scalability but sacrifice optimality altogether, resulting in suboptimal execution schedules. The root cause of these inefficiencies in the state of the art is the mismatch in abstraction. This paper introduces Bonsai, a framework that tackles this scalability-granularity tension. At the heart of Bonsai is a novel abstraction of operator segmentation that breaks the computation graph into flexible, variable-sized units to enable a lightweight yet effective segment-based ILP formulation. By having segments, Bonsai collapses the search space and prunes redundant solutions that stall existing solvers. This abstraction enables Bonsai to maintain a holistic view of the entire model, ensuring that no optimization opportunity is lost while reducing the number of decision variables by orders of magnitude. The evaluation across a diverse set of DNN architectures and models demonstrates that Bonsai scales to real-world models, is up to 10.13× lower solver cost than state-of-the-art ILP solvers, and delivers up to 11
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- TSPLIT: Fine-grained GPU Memory Management for Efficient DNN Training via Tensor SplittingXiaonan Nie, Xupeng Miao, Zhi Yang, Bin CuiICDE 2022 · 被引用 26 次
- HiRemate: Hierarchical Approach for Efficient Re-materialization of Neural NetworksJulia Gusak, Xunyi Zhao, Théotime Le Hellard, Zhe Li 等ICML 2025
- Bonsai: Gradient-free Graph Condensation for Node ClassificationMridul Gupta, Samyak Jain, Vansh Ramani, Hariprasad Kodamana 等ICLR 2025
- A fast heuristic to optimize time-space tradeoff for large modelsAkifumi Imanishi, Zijian Xu, Masayuki Takagi, Sixue Wang 等NeurIPS 2023 · 被引用 2 次
- MODeL: Memory Optimizations for Deep LearningBenoit Steiner, Mostafa Elhoushi, Jacob Kahn, James HegartyICML 2023 · 被引用 17 次
