The Sparsity-Aware LazyGPU Architecture
Changxi Liu, Miao Yu, Yifan Sun, Trevor E. Carlson
摘要
General-Purpose Graphics Processing Units (GPUs) are essential accelerators in data-parallel applications, including machine learning, and physical simulations. Although GPUs utilize fast wavefront context switching to hide memory access latency, memory continues to be a significant bottleneck, limiting overall performance for many important workloads. Current GPU hardware enhancements focus on issuing memory requests in advance to help solve the memory bandwidth bottleneck and improve GPU performance. However, this approach can still be inefficient, leading to hardware contention and suboptimal resource utilization. Instead of issuing memory requests in advance, we take an alternative view on improving GPU performance: lazily issuing memory requests to eliminate memory requests where either (a) the fetched values are zero or (b) they do not affect the result of the executing workload. Building on these insights, we propose LazyGPU, which integrates lazy execution cores with a Zero Cache to eliminate memory requests when all data required by a wavefront is zero. Moreover, LazyGPU utilizes instructions, including multiplication and multiply-add, to eliminate memory requests whose fetched values do not affect the outcomes of these instructions. For example, LazyGPU achieves a 2.18 × speedup compared with the baseline at 60% weight sparsity for the inference of LLaMA 7B.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- Dual-side Sparse Tensor CoreYang Wang, Chen Zhang, Zhiqiang Xie, Cong Guo 等ISCA 2021 · 被引用 109 次
- Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware PrefetchingSamuel Pakalapati, Biswabandan PandaISCA 2020 · 被引用 97 次
- Accelerating Large Scale Real-Time GNN Inference using Channel PruningHongkuan Zhou, Ajitesh Srivastava, Hanqing Zeng, Rajgopal Kannan 等VLDB 2021 · 被引用 86 次
- GoSPA: An Energy-efficient High-performance Globally Optimized SParse Convolutional Neural Network AcceleratorChunhua Deng, Yang Sui, Siyu Liao, Xuehai Qian 等ISCA 2021 · 被引用 77 次
相关 Paper
- Snake: A Variable-length Chain-based Prefetching for GPUsSaba Mostofi, Hajar Falahati, Negin Mahani, Pejman Lotfi-Kamran 等MICRO 2023 · 被引用 11 次
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 被引用 13 次
- Understand and Accelerate Memory Processing Pipeline for Large Language Model InferenceZifan He, Rui Ma, Yizhou Sun, Jason CongICML 2026
- Morpheus: Extending the Last Level Cache Capacity in GPU Systems Using Idle GPU Core ResourcesSina Darabi, Mohammad Sadrosadati, Negar Akbarzadeh, Joël Lindegger 等MICRO 2022 · 被引用 24 次
- NearFetch: Saving Inter-Module Bandwidth in Many-Chip-Module GPUsXia Zhao, Guangda Zhang, Lu Wang, Shiqing Zhang 等HPCA 2025 · 被引用 3 次
