Rockmate: an Efficient, Fast, Automatic and Generic Tool for Re-materialization in PyTorch
Xunyi Zhao, Théotime Le Hellard, Lionel Eyraud-Dubois, Julia Gusak, Olivier Beaumont
摘要
We propose Rockmate to control the memory requirements when training PyTorch DNN models. Rockmate is an automatic tool that starts from the model code and generates an equivalent model, using a predefined amount of memory for activations, at the cost of a few re-computations. Rockmate automatically detects the structure of computational and data dependencies and rewrites the initial model as a sequence of complex blocks. We show that such a structure is widespread and can be found in many models in the literature (Transformer based models, ResNet, RegNets,...). This structure allows us to solve the problem in a fast and efficient way, using an adaptation of Checkmate (too slow on the whole model but general) at the level of individual blocks and an adaptation of Rotor (fast but limited to sequential models) at the level of the sequence itself. We show through experiments on many models that Rockmate is as fast as Rotor and as efficient as Checkmate, and that it allows in many cases to obtain a significantly lower memory consumption for activations (by a factor of 2 to 5) for a rather negligible overhead (of the order of 10% to 20%). Rockmate is open source and available at https: //github.com/topal-team/rockmate .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SpeedLoader: An I/O efficient scheme for heterogeneous and distributed LLM operationYiqi Zhang, Yang YouNeurIPS 2024 · 被引用 6 次
- A fast heuristic to optimize time-space tradeoff for large modelsAkifumi Imanishi, Zijian Xu, Masayuki Takagi, Sixue Wang 等NeurIPS 2023 · 被引用 2 次
- HiRemate: Hierarchical Approach for Efficient Re-materialization of Neural NetworksJulia Gusak, Xunyi Zhao, Théotime Le Hellard, Zhe Li 等ICML 2025
它引用的顶会 Paper4
- Dynamic Tensor RematerializationMarisa Kirisame, Steven Lyubomirsky, Altan Haan, Jennifer Brennan 等ICLR 2021 · 被引用 115 次
- Varuna: scalable, low-cost training of massive deep learning modelsSanjith Athlur, Nitika Saran, Muthian Sivathanu, Ramachandran Ramjee 等EuroSys 2022 · 被引用 81 次
- Efficient Combination of Rematerialization and Offloading for Training DNNsOlivier Beaumont, Lionel Eyraud-Dubois, Alena ShilovaNeurIPS 2021 · 被引用 69 次
- POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and PagingShishir G. Patil, Paras Jain, Prabal Dutta, Ion Stoica 等ICML 2022 · 被引用 52 次
相关 Paper
- Memory Optimization for Deep NetworksAashaka Shah, Chao-Yuan Wu, Jayashree Mohan, Vijay Chidambaram 等ICLR 2021 · 被引用 29 次
- Occamy: Memory-efficient GPU Compiler for DNN InferenceJaeho Lee, Shinnung Jeong, Seungbin Song, Kunwoo Kim 等DAC 2023 · 被引用 3 次
- T-Control: An Efficient Dynamic Tensor Rematerialization System for DNN TrainingZehua Wang, Junmin Xiao, Xiaochuan Deng, Huibing Wang 等ASPLOS 2026 · 被引用 2 次
- EXACT: Scalable Graph Neural Networks Training via Extreme Activation CompressionZirui Liu, Kaixiong Zhou, Fan Yang, Li Li 等ICLR 2022 · 被引用 72 次
- Sparse Weight Activation TrainingMd Aamir Raihan, Tor M. AamodtNeurIPS 2020 · 被引用 83 次
