η-LSTM: Co-Designing Highly-Efficient Large LSTM Training via Exploiting Memory-Saving and Architectural Design Opportunities
Xingyao Zhang, Haojun Xia, Donglin Zhuang, Hao Sun, Xin Fu, Michael B. Taylor, Shuaiwen Leon Song
摘要
Recently, the recurrent neural network, or its most popular type—the Long Short Term Memory (LSTM) network— has achieved great success in a broad spectrum of real-world application domains, such as autonomous driving, natural language processing, sentiment analysis, and epidemiology. Due to the complex features of the real-world tasks, current LSTM models become increasingly bigger and more complicated for enhancing the learning ability and prediction accuracy. However, through our in-depth characterization on the state-of-the-art general-purpose deep-learning accelerators, we observe that the LSTM training execution grows inefficient in terms of storage, performance, and energy consumption, under an increasing model size. With further algorithmic and architectural analysis, we identify the root cause for large LSTM training inefficiency: massive intermediate variables. To enable a highly-efficient LSTM training solution for the ever-growing model size, we exploit some unique memory-saving and performance improvement opportunities from the LSTM training procedure, and leverage them to propose the first cross-stack training solution, η-LSTM, for large LSTM models. η-LSTM comprises both software-level and hardware-level innovations that effectively lower the memory footprint upper-bound and excessive data movements during large LSTM training, while also drastically improving training performance and energy efficiency. Experimental results on six real-world large LSTM training benchmarks demonstrate that η-LSTM reduces the required memory footprint by an average of 57.5% (up to 75.8%) and brings down the data movements for weight matrices, activation data, and intermediate variables by 40.9%, 32.9%, and 80.0%, respectively. Furthermore, it outperforms the state-of-the-art GPU implementation for LSTM training by an average of 3.99× (up to 5.73×) on performance and 2.75× (up to 4.25) on energy. We hope this work can shed some light on how to design high logic utilization for future NPUs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Ditto: Accelerating Diffusion Model via Temporal Value SimilaritySungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park 等HPCA 2025 · 被引用 9 次
- TaGNN: An Efficient Topology-aware Accelerator for High-performance Dynamic Graph Neural NetworkHui Yu, Yu Zhang, Ligang He, Bing Peng 等SC 2025 · 被引用 2 次
它引用的顶会 Paper5
- TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network TrainingMostafa Mahmoud, Isak Edo, Ali Hadi Zadeh, Omar Mohamed Awad 等MICRO 2020 · 被引用 78 次
- Procrustes: a Dataflow and Accelerator for Sparse Deep Neural Network TrainingDingqing Yang, Amin Ghasemazar, Xiaowei Ren, Maximilian Golub 等MICRO 2020 · 被引用 63 次
- SmartExchange: Trading Higher-cost Memory Storage/Access for Lower-cost ComputationYang Zhao, Xiaohan Chen, Yue Wang, Chaojian Li 等ISCA 2020 · 被引用 44 次
- Echo: Compiler-based GPU Memory Footprint Reduction for LSTM RNN TrainingBojian Zheng, Nandita Vijaykumar, Gennady PekhimenkoISCA 2020 · 被引用 34 次
- Enabling Highly Efficient Capsule Networks Processing Through A PIM-Based Architecture DesignXingyao Zhang, Shuaiwen Leon Song, Chenhao Xie, Jing Wang 等HPCA 2020 · 被引用 22 次
相关 Paper
- Vision-LSTM: xLSTM as Generic Vision BackboneBenedikt Alkin, Maximilian Beck, Korbinian Pöppel, Sepp Hochreiter 等ICLR 2025 · 被引用 20 次
- Parallelizing Legendre Memory Unit TrainingNarsimha Reddy Chilkuri, Chris EliasmithICML 2021 · 被引用 47 次
- Unlocking the Power of LSTM for Long Term Time Series ForecastingYaxuan Kong, Zepu Wang, Yuqi Nie, Tian Zhou 等AAAI 2025 · 被引用 92 次
- xLSTM: Extended Long Short-Term MemoryMaximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer 等NeurIPS 2024 · 被引用 703 次
- MomentumRNN: Integrating Momentum into Recurrent Neural NetworksTan M. Nguyen, Richard G. Baraniuk, Andrea L. Bertozzi, Stanley J. Osher 等NeurIPS 2020 · 被引用 32 次
