Improving Data Reuse in NPU On-chip Memory with Interleaved Gradient Order for DNN Training
Jungwoo Kim, Seonjin Na, Sanghyeon Lee, Sunho Lee, Jaehyuk Huh
摘要
During training tasks for machine learning models with neural processing units (NPUs), the most time-consuming part is the backward pass, which incurs significant overheads due to off-chip memory accesses. For NPUs, to mitigate the long latency and limited bandwidth of such off-chip DRAM accesses, the software-managed onchip scratchpad memory (SPM) plays a crucial role. As the backward pass computation must be optimized to improve the effectiveness of SPM, this study identifies a new data reuse pattern specific to the backward computation. The backward pass includes independent input and weight gradient computations sharing the same output gradient in each layer. Conventional sequential processing does not exploit the potential inter-operation data reuse opportunity within SPM. With this new opportunity of data reuse in the backward pass, this study proposes a novel data flow transformation scheme called interleaved gradient order, consisting of three techniques to enhance the utilization of NPU scratchpad memory. The first technique shuffles the input and weight gradient computations by interleaving two operations into a single fused operation to reduce redundant output gradient accesses. The second technique adjusts the tile access order for the interleaved gradient computations to maximize the potential data locality. However, since the best order is not fixed for all tensors, we propose a selection algorithm to find the most suitable order based on the tensor dimensions. The final technique further improves data reuse chances by using the best partitioning and mapping scheme for two gradient computations for single-core and multi-core NPUs. The simulation-based evaluation with singlecore edge and server NPUs shows that the combined techniques can improve performance by 29.3% and 14.5% for edge and server NPUs respectively. Furthermore, with a quad-core server NPU, the proposed techniques reduce the execution time by 23.7%. * Seonjin Na is currently with Georgia Institute of Technology.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella 等HPCA 2020 · 被引用 490 次
- Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationHasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali 等DAC 2021 · 被引用 325 次
相关 Paper
- EDA: Energy-Efficient Inter-Layer Model Compilation for Edge DNN Inference AccelerationBo Ren Pao, I-Chia Chen, En-Hao Chang, Tsung Tai YehHPCA 2025 · 被引用 1 次
- Register Tiling for Unstructured Sparsity in Neural Network InferenceLucas Wilkinson, Kazem Cheshmi, Maryam Mehri DehnaviPLDI 2023 · 被引用 17 次
- Out-of-order backprop: an effective scheduling technique for deep learningHyungjun Oh, Junyeol Lee, HyeongJu Kim, Jiwon SeoEuroSys 2022 · 被引用 14 次
- A Pragmatic Approach to On-device Incremental Learning System with Selective Weight UpdatesJaekang Shin, Seungkyu Choi, Yeongjae Choi, Lee-Sup KimDAC 2020 · 被引用 7 次
- Dataflow Mirroring: Architectural Support for Highly Efficient Fine-Grained Spatial Multitasking on Systolic-Array NPUsJounghoo Lee, Jinwoo Choi, Jaeyeon Kim, Jinho Lee 等DAC 2021 · 被引用 39 次
