η-LSTM: Co-Designing Highly-Efficient Large LSTM Training via Exploiting Memory-Saving and Architectural Design Opportunities
Xingyao Zhang, Haojun Xia, Donglin Zhuang, Hao Sun, Xin Fu, Michael B. Taylor, Shuaiwen Leon Song
Abstract
Recently, the recurrent neural network, or its most popular type—the Long Short Term Memory (LSTM) network— has achieved great success in a broad spectrum of real-world application domains, such as autonomous driving, natural language processing, sentiment analysis, and epidemiology. Due to the complex features of the real-world tasks, current LSTM models become increasingly bigger and more complicated for enhancing the learning ability and prediction accuracy. However, through our in-depth characterization on the state-of-the-art general-purpose deep-learning accelerators, we observe that the LSTM training execution grows inefficient in terms of storage, performance, and energy consumption, under an increasing model size. With further algorithmic and architectural analysis, we identify the root cause for large LSTM training inefficiency: massive intermediate variables. To enable a highly-efficient LSTM training solution for the ever-growing model size, we exploit some unique memory-saving and performance improvement opportunities from the LSTM training procedure, and leverage them to propose the first cross-stack training solution, η-LSTM, for large LSTM models. η-LSTM comprises both software-level and hardware-level innovations that effectively lower the memory footprint upper-bound and excessive data movements during large LSTM training, while also drastically improving training performance and energy efficiency. Experimental results on six real-world large LSTM training benchmarks demonstrate that η-LSTM reduces the required memory footprint by an average of 57.5% (up to 75.8%) and brings down the data movements for weight matrices, activation data, and intermediate variables by 40.9%, 32.9%, and 80.0%, respectively. Furthermore, it outperforms the state-of-the-art GPU implementation for LSTM training by an average of 3.99× (up to 5.73×) on performance and 2.75× (up to 4.25) on energy. We hope this work can shed some light on how to design high logic utilization for future NPUs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 873da0ef-6f8b-4837-818c-2310d581de2cCited by top-tier papers2
- Ditto: Accelerating Diffusion Model via Temporal Value SimilaritySungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park et al.HPCA 2025 · 9 citations
- TaGNN: An Efficient Topology-aware Accelerator for High-performance Dynamic Graph Neural NetworkHui Yu, Yu Zhang, Ligang He, Bing Peng et al.SC 2025 · 2 citations
Builds on5
- TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network TrainingMostafa Mahmoud, Isak Edo, Ali Hadi Zadeh, Omar Mohamed Awad et al.MICRO 2020 · 78 citations
- Procrustes: a Dataflow and Accelerator for Sparse Deep Neural Network TrainingDingqing Yang, Amin Ghasemazar, Xiaowei Ren, Maximilian Golub et al.MICRO 2020 · 63 citations
- SmartExchange: Trading Higher-cost Memory Storage/Access for Lower-cost ComputationYang Zhao, Xiaohan Chen, Yue Wang, Chaojian Li et al.ISCA 2020 · 44 citations
- Echo: Compiler-based GPU Memory Footprint Reduction for LSTM RNN TrainingBojian Zheng, Nandita Vijaykumar, Gennady PekhimenkoISCA 2020 · 34 citations
- Enabling Highly Efficient Capsule Networks Processing Through A PIM-Based Architecture DesignXingyao Zhang, Shuaiwen Leon Song, Chenhao Xie, Jing Wang et al.HPCA 2020 · 22 citations
Related papers
- Vision-LSTM: xLSTM as Generic Vision BackboneBenedikt Alkin, Maximilian Beck, Korbinian Pöppel, Sepp Hochreiter et al.ICLR 2025 · 20 citations
- Parallelizing Legendre Memory Unit TrainingNarsimha Reddy Chilkuri, Chris EliasmithICML 2021 · 47 citations
- Unlocking the Power of LSTM for Long Term Time Series ForecastingYaxuan Kong, Zepu Wang, Yuqi Nie, Tian Zhou et al.AAAI 2025 · 92 citations
- xLSTM: Extended Long Short-Term MemoryMaximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer et al.NeurIPS 2024 · 703 citations
- MomentumRNN: Integrating Momentum into Recurrent Neural NetworksTan M. Nguyen, Richard G. Baraniuk, Andrea L. Bertozzi, Stanley J. Osher et al.NeurIPS 2020 · 32 citations
