Efficient Combination of Rematerialization and Offloading for Training DNNs
Olivier Beaumont, Lionel Eyraud-Dubois, Alena Shilova
Abstract
Rematerialization and offloading are two well known strategies to save memory during the training phase of deep neural networks, allowing data scientists to consider larger models, batch sizes or higher resolution data. Rematerialization trades memory for computation time, whereas Offloading trades memory for data movements. As these two resources are independent, it is appealing to consider the simultaneous combination of both strategies to save even more memory. We precisely model the costs and constraints corresponding to Deep Learning frameworks such as PyTorch or Tensorflow, we propose optimal algorithms to find a valid sequence of memory-constrained operations and finally, we evaluate the performance of proposed algorithms on realistic networks and computation platforms. Our experiments show that the possibility to offload can remove one third of the overhead of rematerialization, and that together they can reduce the memory used for activations by a factor 4 to 6, with an overhead below 20%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f6b83d0-7b49-4c2b-a4c2-ec94c222c860Cited by top-tier papers20
- POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and PagingShishir G. Patil, Paras Jain, Prabal Dutta, Ion Stoica et al.ICML 2022 · 52 citations
- GACT: Activation Compressed Training for Generic Network ArchitecturesXiaoxuan Liu, Lianmin Zheng, Dequan Wang, Yukuo Cen et al.ICML 2022 · 43 citations
- Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hybrid ParallelismTailing Yuan, Yuliang Liu, Xucheng Ye, Shenglong Zhang et al.USENIX ATC 2024 · 34 citations
- GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory StitchingCong Guo, Rui Zhang, Jiale Xu, Jingwen Leng et al.ASPLOS 2024 · 30 citations
- AdaPipe: Optimizing Pipeline Parallelism with Adaptive Recomputation and PartitioningZhenbo Sun, Huanqi Cao, Yuanwei Wang, Guanyu Feng et al.ASPLOS 2024 · 28 citations
Builds on5
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Memory-Efficient Pipeline-Parallel DNN TrainingDeepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen et al.ICML 2021 · 283 citations
- SwapAdvisor: Pushing Deep Learning Beyond the GPU Memory Limit via Smart SwappingChien-Chin Huang, Gu Jin, Jinyang LiASPLOS 2020 · 161 citations
- Dynamic Tensor RematerializationMarisa Kirisame, Steven Lyubomirsky, Altan Haan, Jennifer Brennan et al.ICLR 2021 · 115 citations
- Optimal Gradient Checkpoint Search for Arbitrary Computation GraphsJianwei Feng, Dong HuangCVPR 2021
Related papers
- Coop: Memory is not a CommodityJianhao Zhang, Shihan Ma, Peihong Liu, Jinhui YuanNeurIPS 2023 · 11 citations
- HiRemate: Hierarchical Approach for Efficient Re-materialization of Neural NetworksJulia Gusak, Xunyi Zhao, Théotime Le Hellard, Zhe Li et al.ICML 2025
- Memory Optimization for Deep NetworksAashaka Shah, Chao-Yuan Wu, Jayashree Mohan, Vijay Chidambaram et al.ICLR 2021 · 29 citations
- T-Control: An Efficient Dynamic Tensor Rematerialization System for DNN TrainingZehua Wang, Junmin Xiao, Xiaochuan Deng, Huibing Wang et al.ASPLOS 2026 · 2 citations
- SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model TrainingKun Wu, Jeongmin Brian Park, Xiaofan Zhang, Mert Hidayetoglu et al.DAC 2025 · 3 citations
