Tempo: Accelerating Transformer-Based Model Training through Memory Footprint Reduction
Muralidhar Andoorveedu, Zhanda Zhu, Bojian Zheng, Gennady Pekhimenko
Abstract
Training deep learning models can be computationally expensive. Prior works have shown that increasing the batch size can potentially lead to better overall throughput. However, the batch size is frequently limited by the accelerator memory capacity due to the activations/feature maps stored for the training backward pass, as larger batch sizes require larger feature maps to be stored. Transformer-based models, which have recently seen a surge in popularity due to their good performance and applicability to a variety of tasks, have a similar problem. To remedy this issue, we propose Tempo, a new approach to efficiently use accelerator (e.g., GPU) memory resources for training Transformer-based models. Our approach provides drop-in replacements for the GELU, LayerNorm, and Attention layers, reducing the memory usage and ultimately leading to more efficient training. We implement Tempo and evaluate the throughput, memory usage, and accuracy/loss on the BERT Large pre-training task. We demonstrate that Tempo enables up to 2x higher batch sizes and 16% higher training throughput over the state-of-the-art baseline. We also evaluate Tempo on GPT2 and RoBERTa models, showing 19% and 26% speedup over the baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d67852ad-0f7b-43b4-bac7-5dfd468bbb7dCited by top-tier papers3
- Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language ModelZirui Liu, Guanchu Wang, Shaochen Zhong, Zhaozhuo Xu et al.NeurIPS 2023 · 22 citations
- Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-OptimizationZhanda Zhu, Christina Giannoula, Muralidhar Andoorveedu, Qidong Su et al.EuroSys 2025 · 8 citations
- HawkEye: Statically and Accurately Profiling the Communication Cost of Models in Multi-party LearningWenqiang Ruan, Xin Lin, Ruisheng Zhou, Guopeng Lin et al.USENIX Security 2025
Builds on8
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith et al.SC 2021 · 254 citations
- Capuchin: Tensor-based GPU Memory Management for Deep LearningXuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin et al.ASPLOS 2020 · 143 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
- Dynamic Tensor RematerializationMarisa Kirisame, Steven Lyubomirsky, Altan Haan, Jennifer Brennan et al.ICLR 2021 · 115 citations
Related papers
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 126 citations
- LightSeq2: Accelerated Training for Transformer-Based Models on GPUsXiaohui Wang, Yang Wei, Ying Xiong, Guyue Huang et al.SC 2022 · 25 citations
- H3T: Efficient Integration of Memory Optimization and Parallelism for Large-scale Transformer TrainingYuzhong Wang, Xu Han, Weilin Zhao, Guoyang Zeng et al.NeurIPS 2023 · 2 citations
- CoMERA: Computing- and Memory-Efficient Training via Rank-Adaptive Tensor OptimizationZi Yang, Ziyue Liu, Samridhi Choudhary, Xinfeng Xie et al.NeurIPS 2024 · 18 citations
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu et al.ACL 2022
