Tempo: Accelerating Transformer-Based Model Training through Memory Footprint Reduction
Muralidhar Andoorveedu, Zhanda Zhu, Bojian Zheng, Gennady Pekhimenko
摘要
Training deep learning models can be computationally expensive. Prior works have shown that increasing the batch size can potentially lead to better overall throughput. However, the batch size is frequently limited by the accelerator memory capacity due to the activations/feature maps stored for the training backward pass, as larger batch sizes require larger feature maps to be stored. Transformer-based models, which have recently seen a surge in popularity due to their good performance and applicability to a variety of tasks, have a similar problem. To remedy this issue, we propose Tempo, a new approach to efficiently use accelerator (e.g., GPU) memory resources for training Transformer-based models. Our approach provides drop-in replacements for the GELU, LayerNorm, and Attention layers, reducing the memory usage and ultimately leading to more efficient training. We implement Tempo and evaluate the throughput, memory usage, and accuracy/loss on the BERT Large pre-training task. We demonstrate that Tempo enables up to 2x higher batch sizes and 16% higher training throughput over the state-of-the-art baseline. We also evaluate Tempo on GPT2 and RoBERTa models, showing 19% and 26% speedup over the baseline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language ModelZirui Liu, Guanchu Wang, Shaochen Zhong, Zhaozhuo Xu 等NeurIPS 2023 · 被引用 22 次
- Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-OptimizationZhanda Zhu, Christina Giannoula, Muralidhar Andoorveedu, Qidong Su 等EuroSys 2025 · 被引用 8 次
- HawkEye: Statically and Accurately Profiling the Communication Cost of Models in Multi-party LearningWenqiang Ruan, Xin Lin, Ruisheng Zhou, Guopeng Lin 等USENIX Security 2025
它引用的顶会 Paper8
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith 等SC 2021 · 被引用 254 次
- Capuchin: Tensor-based GPU Memory Management for Deep LearningXuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin 等ASPLOS 2020 · 被引用 143 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
- Dynamic Tensor RematerializationMarisa Kirisame, Steven Lyubomirsky, Altan Haan, Jennifer Brennan 等ICLR 2021 · 被引用 115 次
相关 Paper
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 被引用 126 次
- LightSeq2: Accelerated Training for Transformer-Based Models on GPUsXiaohui Wang, Yang Wei, Ying Xiong, Guyue Huang 等SC 2022 · 被引用 25 次
- H3T: Efficient Integration of Memory Optimization and Parallelism for Large-scale Transformer TrainingYuzhong Wang, Xu Han, Weilin Zhao, Guoyang Zeng 等NeurIPS 2023 · 被引用 2 次
- CoMERA: Computing- and Memory-Efficient Training via Rank-Adaptive Tensor OptimizationZi Yang, Ziyue Liu, Samridhi Choudhary, Xinfeng Xie 等NeurIPS 2024 · 被引用 18 次
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu 等ACL 2022
