A fast heuristic to optimize time-space tradeoff for large models
Akifumi Imanishi, Zijian Xu, Masayuki Takagi, Sixue Wang, Emilio Castillo
摘要
Training large-scale neural networks is heavily constrained by GPU memory. In order to circumvent this limitation, gradient checkpointing, or recomputation is a powerful technique. There is active research in this area with methods such as Checkmake [19] or Moccasin [3] . However, both Checkmate and Moccasin rely on mixed integer linear programming or constraint programming, resulting in limited scalability due to their exponentially large search space. This paper proposes a novel algorithm for recomputation (FastSA) based on a simulated annealing heuristic that achieves comparable or even better solutions than state-of-the-art alternatives. FastSA can optimize computational graphs with thousands of nodes within 3 to 30 seconds, several orders of magnitude faster than current solutions. We applied FastSA to PyTorch models and verified its effectiveness through popular large vision and text models, including recent language models with the transformer architecture. The results demonstrate significant memory reductions by 73% with extra 18% computational overheads on average. Our experiments demonstrate the practicality and effectiveness of our recomputation algorithm, further highlighting its potential for wide application in various deep learning domains. * Equal contribution 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- Optimal Gradient Checkpoint Search for Arbitrary Computation GraphsJianwei Feng, Dong HuangCVPR 2021
- Memory Optimization for Deep NetworksAashaka Shah, Chao-Yuan Wu, Jayashree Mohan, Vijay Chidambaram 等ICLR 2021 · 被引用 29 次
- Moccasin: Efficient Tensor Rematerialization for Neural NetworksBurak Bartan, Haoming Li, Harris Teague, Christopher Lott 等ICML 2023 · 被引用 3 次
- Bonsai: Efficient and Optimal Automatic Tensor Rematerialization for Memory-Constrained DNN TrainingDat Nguyen, Vasudha Devarakonda, Anxiao Jiang, Khanh NguyenOOPSLA 2026
- Dynamic Tensor RematerializationMarisa Kirisame, Steven Lyubomirsky, Altan Haan, Jennifer Brennan 等ICLR 2021 · 被引用 115 次
