Dynamic Tensor Rematerialization
Marisa Kirisame, Steven Lyubomirsky, Altan Haan, Jennifer Brennan, Mike He, Jared Roesch, Tianqi Chen, Zachary Tatlock
摘要
Checkpointing enables the training of deep learning models under restricted memory budgets by freeing intermediate activations from memory and recomputing them on demand. Current checkpointing techniques statically plan these recomputations offline and assume static computation graphs. We demonstrate that a simple online algorithm can achieve comparable performance by introducing Dynamic Tensor Rematerialization (DTR), a greedy online algorithm for checkpointing that is extensible and general, is parameterized by eviction policy, and supports dynamic models. We prove that DTR can train an N -layer linear feedforward network on an Ω( √ N ) memory budget with only O(N ) tensor operations. DTR closely matches the performance of optimal static checkpointing in simulated experiments. We incorporate a DTR prototype into PyTorch merely by interposing on tensor allocations and operator calls and collecting lightweight metadata on tensors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless ThreadsJohn Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng 等OSDI 2021 · 被引用 175 次
- Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and ParallelizationColin Unger, Zhihao Jia, Wei Wu, Sina Lin 等OSDI 2022 · 被引用 105 次
- ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed TrainingJianfei Chen, Lianmin Zheng, Zhewei Yao, Dequan Wang 等ICML 2021 · 被引用 93 次
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningLianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang 等OSDI 2022 · 被引用 75 次
- Efficient Combination of Rematerialization and Offloading for Training DNNsOlivier Beaumont, Lionel Eyraud-Dubois, Alena ShilovaNeurIPS 2021 · 被引用 69 次
它引用的顶会 Paper3
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- SwapAdvisor: Pushing Deep Learning Beyond the GPU Memory Limit via Smart SwappingChien-Chin Huang, Gu Jin, Jinyang LiASPLOS 2020 · 被引用 161 次
- Capuchin: Tensor-based GPU Memory Management for Deep LearningXuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin 等ASPLOS 2020 · 被引用 143 次
相关 Paper
- T-Control: An Efficient Dynamic Tensor Rematerialization System for DNN TrainingZehua Wang, Junmin Xiao, Xiaochuan Deng, Huibing Wang 等ASPLOS 2026 · 被引用 2 次
- Coop: Memory is not a CommodityJianhao Zhang, Shihan Ma, Peihong Liu, Jinhui YuanNeurIPS 2023 · 被引用 11 次
- Memory Optimization for Deep NetworksAashaka Shah, Chao-Yuan Wu, Jayashree Mohan, Vijay Chidambaram 等ICLR 2021 · 被引用 29 次
- HiRemate: Hierarchical Approach for Efficient Re-materialization of Neural NetworksJulia Gusak, Xunyi Zhao, Théotime Le Hellard, Zhe Li 等ICML 2025
- CheckFreq: Frequent, Fine-Grained DNN CheckpointingJayashree Mohan, Amar Phanishayee, Vijay ChidambaramFAST 2021 · 被引用 175 次
