Echo: Compiler-based GPU Memory Footprint Reduction for LSTM RNN Training
Bojian Zheng, Nandita Vijaykumar, Gennady Pekhimenko
摘要
The Long-Short-Term-Memory Recurrent Neural Networks (LSTM RNNs) are a popular class of machine learning models for analyzing sequential data. Their training on modern GPUs, however, is limited by the GPU memory capacity. Our profiling results of the LSTM RNN-based Neural Machine Translation (NMT) model reveal that feature maps of the attention and RNN layers form the memory bottleneck, and runtime is unevenly distributed across different layers when training on GPUs. Based on these two observations, we propose to recompute the feature maps of the attention and RNN layers rather than stashing them persistently in the GPU memory. While the idea of feature map recomputation has been considered before, existing solutions fail to deliver satisfactory footprint reduction, as they do not address two key challenges. For each feature map recomputation to be efficient, its effect on (1) the total memory footprint, and (2) the total execution time has to be carefully estimated. To this end, we propose Echo, a new compiler-based optimization scheme that addresses the first challenge with a practical mechanism that estimates the memory benefits of recomputation over the entire computation graph, and the second challenge by non-conservatively estimating the recomputation runtime overhead leveraging layer specifics. Echo reduces the GPU memory footprint automatically and transparently without any changes required to the training source code, and is effective for models beyond LSTM RNNs. We evaluate Echo on numerous state-of-the-art machine learning workloads, including NMT, DeepSpeech2, Transformer, and ResNet, on real systems with modern GPUs and observe footprint reduction ratios of 1. 89x on average and 3. 13x maximum. Such reduction can be converted into faster training with a larger batch size, savings in GPU energy consumption (e.g., training with one GPU as fast as with four), and/or an increase in the maximum number of layers under the same GPU memory budget. Echo is open-sourced as a part of the MXNet 2.0 framework.11https://issues.apache.org/jirdprojects/MXNET/issues/MXNET-1450
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- AutoFL: Enabling Heterogeneity-Aware Energy Efficient Federated LearningYoung Geun Kim, Carole-Jean WuMICRO 2021 · 被引用 84 次
- Zico: Efficient GPU Memory Sharing for Concurrent DNN TrainingGangmuk Lim, Jeongseob Ahn, Wencong Xiao, Youngjin Kwon 等USENIX ATC 2021 · 被引用 65 次
- FlashNeuron: SSD-Enabled Large-Batch Training of Very Deep Neural NetworksJonghyun Bae, Jongsung Lee, Yunho Jin, Sam Son 等FAST 2021 · 被引用 64 次
- Hidet: Task-Mapping Programming Paradigm for Deep Learning Tensor ProgramsYaoyao Ding, Cody Hao Yu, Bojian Zheng, Yizhi Liu 等ASPLOS 2023 · 被引用 27 次
- MODeL: Memory Optimizations for Deep LearningBenoit Steiner, Mostafa Elhoushi, Jacob Kahn, James HegartyICML 2023 · 被引用 17 次
相关 Paper
- η-LSTM: Co-Designing Highly-Efficient Large LSTM Training via Exploiting Memory-Saving and Architectural Design OpportunitiesXingyao Zhang, Haojun Xia, Donglin Zhuang, Hao Sun 等ISCA 2021 · 被引用 7 次
- Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient TrainingAnup Sarma, Sonali Singh, Huaipan Jiang, Rui Zhang 等NeurIPS 2021 · 被引用 1 次
- Parallelizing Legendre Memory Unit TrainingNarsimha Reddy Chilkuri, Chris EliasmithICML 2021 · 被引用 47 次
- Out of the Memory Barrier: A Highly Memory-Efficient Training System for LLMs with Million-Token ContextsWenhao Li, Daohai Yu, Gen Luo, Yuxin Zhang 等ICLR 2026 · 被引用 5 次
- MEMO: Fine-grained Tensor Management For Ultra-long Context LLM TrainingPinxue Zhao, Hailin Zhang, Fangcheng Fu, Xiaonan Nie 等SIGMOD 2025 · 被引用 4 次
