VENOM: A Vectorized N: M Format for Unleashing the Power of Sparse Tensor Cores
Roberto L. Castro, Andrei Ivanov, Diego Andrade, Tal Ben-Nun, Basilio B. Fraguela, Torsten Hoefler
摘要
The increasing success and scaling of Deep Learning models demands higher computational efficiency and power. Sparsification can lead to both smaller models as well as higher compute efficiency, and accelerated hardware is becoming available. However, exploiting it efficiently requires kernel implementations, pruning algorithms, and storage formats, to utilize hardware support of specialized sparse vector units. An example of those are the NVIDIA's Sparse Tensor Cores (SPTCs), which promise a 2× speedup. However, SPTCs only support the 2:4 format, limiting achievable sparsity ratios to 50%. We present the V:N:M format, which enables the execution of arbitrary N:M ratios on SPTCs. To efficiently exploit the resulting format, we propose Spatha, a high-performance sparse-library for DL routines. We show that Spatha achieves up to 37× speedup over cuBLAS. We also demonstrate a second-order pruning technique that enables sparsification to high sparsity ratios with V:N:M and little to no loss in accuracy in modern transformers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust AdaptationMahdi Nikdan, Soroush Tabesh, Elvir Crncevic, Dan AlistarhICML 2024 · 被引用 53 次
- High Performance Unstructured SpMM Computation Using Tensor CoresPatrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta 等SC 2024 · 被引用 15 次
- DELTA4: Sparse Matrix-Vector Multiplication for Low SparsityVladimír Macko, Vladimír BožaICML 2026 · 被引用 9 次
- GeneralSparse: Bridging the Gap in SpMM for Pruned Large Language Model Inference on GPUsYaoyu Wang, Xiao Guo, Junmin Xiao, De Chen 等USENIX ATC 2025 · 被引用 5 次
- Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor CoresChenpeng Wu, Qiqi Gu, Heng Shi, Jianguo Yao 等EuroSys 2025 · 被引用 5 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 被引用 656 次
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 被引用 170 次
- PLATON: Pruning Large Transformer Models with Upper Confidence Bound of Weight ImportanceQingru Zhang, Simiao Zuo, Chen Liang, Alexander Bukharin 等ICML 2022 · 被引用 107 次
相关 Paper
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou 等PPoPP 2025 · 被引用 18 次
- Efficient Quantized Sparse Matrix Operations on Tensor CoresShigang Li, Kazuki Osawa, Torsten HoeflerSC 2022 · 被引用 27 次
- SPLAT: A Framework for Optimised GPU Code-Generation for SParse reguLar ATtentionAhan Gupta, Yueming Yuan, Devansh Jain, Yuhao Ge 等OOPSLA 2025 · 被引用 3 次
- Bridging the Gap between Unstructured SpMM and Structured Sparse Tensor CoresYukang Dong, Ziyuan Shen, Wenbin Jiang, Zhenghang Liu 等SC 2025 · 被引用 4 次
- Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor CoresKaige Zhang, Hailong Yang, Xin You, Tianyu Feng 等PPoPP 2026
