ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor Cores
Yuetao Chen, Kun Li, Yuhao Wang, Donglin Bai, Lei Wang, Lingxiao Ma, Liang Yuan, Yunquan Zhang, Ting Cao, Mao Yang
摘要
Tensor Core Unit (TCU) is increasingly integrated into modern high-performance processors to enhance matrix multiplication performance. However, constrained to its overspecification, its potential for improving other critical scientific operations like stencil computations remains untapped.
This paper presents ConvStencil 1 , a novel stencil computing system designed to efficiently transform stencil computation to matrix multiplication on Tensor Cores. We first develop a performance model for ConvStencil to guide algorithm design and optimization on TCUs. Based on this model, we propose three techniques: (1) Memory-efficient Layout Transformation using the stencil2row method; (2) * Work done during an internship at Microsoft Research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- AmgT: Algebraic Multigrid Solver on Tensor CoresYuechen Lu, Lijie Zeng, Tengcheng Wang, Xu Fu 等SC 2024 · 被引用 17 次
- LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor CoresYiwei Zhang, Kun Li, Liang Yuan, Jiawen Cheng 等SC 2024 · 被引用 13 次
- FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Tensor Core UnitsHaozhi Han, Kun Li, Wei Cui, Donglin Bai 等PPoPP 2025 · 被引用 7 次
- HStencil: Matrix-Vector Stencil Computation with Interleaved Outer Product and MLAHan Huang, Jiabin Xie, Guangnan Feng, Xianwei Zhang 等SC 2025 · 被引用 5 次
- SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity TransformationQi Li, Kun Li, Haozhi Han, Liang Yuan 等SC 2025 · 被引用 3 次
它引用的顶会 Paper6
- AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstractionSize Zheng, Renze Chen, Anjiang Wei, Yicheng Jin 等ISCA 2022 · 被引用 63 次
- Tacker: Tensor-CUDA Core Kernel Fusion for Improving the GPU Utilization while Ensuring QoSHan Zhao, Weihao Cui, Quan Chen, Youtao Zhang 等HPCA 2022 · 被引用 42 次
- Improving communication by optimizing on-node data movement with data layoutTuowen Zhao, Mary W. Hall, Hans Johansen, Samuel WilliamsPPoPP 2021 · 被引用 19 次
- Productive Performance Engineering for Weather and Climate Modeling with PythonTal Ben-Nun, Linus Groner, Florian Deconinck, Tobias Wicky 等SC 2022 · 被引用 17 次
- Scalable Distributed High-Order Stencil ComputationsMathias Jacquelin, Mauricio Araya-Polo, Jie MengSC 2022 · 被引用 17 次
相关 Paper
- SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided SwappingQiqi Gu, Chenpeng Wu, Heng Shi, Jianguo YaoPPoPP 2026 · 被引用 1 次
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou 等PPoPP 2025 · 被引用 18 次
- A Tensor Marshaling Unit for Sparse Tensor Algebra on General-Purpose ProcessorsMarco Siracusa, Víctor Soria Pardos, Francesco Sgherzi, Joshua Randall 等MICRO 2023 · 被引用 11 次
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu 等PPoPP 2025 · 被引用 18 次
- Optimizing Direct Convolutions on ARM Multi-CoresPengyu Wang, Weiling Yang, Jianbin Fang, Dezun Dong 等SC 2023 · 被引用 6 次
