EGEMM-TC: accelerating scientific computing on tensor cores with extended precision
Boyuan Feng, Yuke Wang, Guoyang Chen, Weifeng Zhang, Yuan Xie, Yufei Ding
摘要
Nvidia Tensor Cores achieve high performance with half-precision matrix inputs tailored towards deep learning workloads. However, this limits the application of Tensor Cores especially in the area of scientific computing with high precision requirements. In this paper, we build Emulated GEMM on Tensor Cores (EGEMM-TC) to extend the usage of Tensor Cores to accelerate scientific computing applications without compromising the precision requirements. First, EGEMM-TC employs an extendable workflow of hardware profiling and operation design to generate a lightweight emulation algorithm on Tensor Cores with extended-precision. Second, EGEMM-TC exploits a set of Tensor Core kernel optimizations to achieve high performance, including the highly-efficient tensorization to exploit the Tensor Core memory architecture and the instruction-level optimizations to coordinate the emulation computation and memory access. Third, EGEMM-TC incorporates a hardware-aware analytic model to offer large flexibility for automatic performance tuning across various scientific computing workloads and input datasets. Extensive evaluations show that EGEMM-TC can achieve on average 3.13× and 11.18× speedup over the cuBLAS kernels and the CUDA-SDK kernels on CUDA Cores, respectively. Our case study on several scientific computing applications further confirms that EGEMM-TC can generalize the usage of Tensor Cores and achieve about 1.8× speedup compared to the hand-tuned, highly-optimized implementations running on CUDA Cores.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- AmgT: Algebraic Multigrid Solver on Tensor CoresYuechen Lu, Lijie Zeng, Tengcheng Wang, Xu Fu 等SC 2024 · 被引用 17 次
- ReFloat: Low-Cost Floating-Point Processing in ReRAM for Accelerating Iterative Linear SolversLinghao Song, Fan Chen, Hai Li, Yiran ChenSC 2023 · 被引用 8 次
- SIMD2: a generalized matrix instruction set for accelerating tensor computation beyond GEMMYunan Zhang, Po-An Tsai, Hung-Wei TsengISCA 2022 · 被引用 6 次
- Misam: Machine Learning Assisted Dataflow Selection in Accelerators for Sparse Matrix MultiplicationSanjali Yadav, Amirmahdi Namjoo, Bahar AsgariMICRO 2025 · 被引用 6 次
- HStencil: Matrix-Vector Stencil Computation with Interleaved Outer Product and MLAHan Huang, Jiabin Xie, Guangnan Feng, Xianwei Zhang 等SC 2025 · 被引用 5 次
相关 Paper
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou 等PPoPP 2025 · 被引用 18 次
- High Performance Unstructured SpMM Computation Using Tensor CoresPatrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta 等SC 2024 · 被引用 15 次
- Efficient Quantized Sparse Matrix Operations on Tensor CoresShigang Li, Kazuki Osawa, Torsten HoeflerSC 2022 · 被引用 27 次
- DTC-SpMM: Bridging the Gap in Accelerating General Sparse Matrix Multiplication with Tensor CoresRuibo Fan, Wei Wang, Xiaowen ChuASPLOS 2024 · 被引用 46 次
- Cooperative Warp Execution in Tensor Core for RISC-V GPGPUAbubakr Nada, Giuseppe Maria Sarda, Erwan LenormandHPCA 2025 · 被引用 3 次
