ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor Cores
Yuetao Chen, Kun Li, Yuhao Wang, Donglin Bai, Lei Wang, Lingxiao Ma, Liang Yuan, Yunquan Zhang, Ting Cao, Mao Yang
Abstract
Tensor Core Unit (TCU) is increasingly integrated into modern high-performance processors to enhance matrix multiplication performance. However, constrained to its overspecification, its potential for improving other critical scientific operations like stencil computations remains untapped.
This paper presents ConvStencil 1 , a novel stencil computing system designed to efficiently transform stencil computation to matrix multiplication on Tensor Cores. We first develop a performance model for ConvStencil to guide algorithm design and optimization on TCUs. Based on this model, we propose three techniques: (1) Memory-efficient Layout Transformation using the stencil2row method; (2) * Work done during an internship at Microsoft Research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6608b360-e5c6-4309-b7fb-06d626554281Cited by top-tier papers9
- AmgT: Algebraic Multigrid Solver on Tensor CoresYuechen Lu, Lijie Zeng, Tengcheng Wang, Xu Fu et al.SC 2024 · 17 citations
- LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor CoresYiwei Zhang, Kun Li, Liang Yuan, Jiawen Cheng et al.SC 2024 · 13 citations
- FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Tensor Core UnitsHaozhi Han, Kun Li, Wei Cui, Donglin Bai et al.PPoPP 2025 · 7 citations
- HStencil: Matrix-Vector Stencil Computation with Interleaved Outer Product and MLAHan Huang, Jiabin Xie, Guangnan Feng, Xianwei Zhang et al.SC 2025 · 5 citations
- SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity TransformationQi Li, Kun Li, Haozhi Han, Liang Yuan et al.SC 2025 · 3 citations
Builds on6
- AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstractionSize Zheng, Renze Chen, Anjiang Wei, Yicheng Jin et al.ISCA 2022 · 63 citations
- Tacker: Tensor-CUDA Core Kernel Fusion for Improving the GPU Utilization while Ensuring QoSHan Zhao, Weihao Cui, Quan Chen, Youtao Zhang et al.HPCA 2022 · 42 citations
- Improving communication by optimizing on-node data movement with data layoutTuowen Zhao, Mary W. Hall, Hans Johansen, Samuel WilliamsPPoPP 2021 · 19 citations
- Productive Performance Engineering for Weather and Climate Modeling with PythonTal Ben-Nun, Linus Groner, Florian Deconinck, Tobias Wicky et al.SC 2022 · 17 citations
- Scalable Distributed High-Order Stencil ComputationsMathias Jacquelin, Mauricio Araya-Polo, Jie MengSC 2022 · 17 citations
Related papers
- SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided SwappingQiqi Gu, Chenpeng Wu, Heng Shi, Jianguo YaoPPoPP 2026 · 1 citation
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou et al.PPoPP 2025 · 18 citations
- A Tensor Marshaling Unit for Sparse Tensor Algebra on General-Purpose ProcessorsMarco Siracusa, Víctor Soria Pardos, Francesco Sgherzi, Joshua Randall et al.MICRO 2023 · 11 citations
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu et al.PPoPP 2025 · 18 citations
- Optimizing Direct Convolutions on ARM Multi-CoresPengyu Wang, Weiling Yang, Jianbin Fang, Dezun Dong et al.SC 2023 · 6 citations
