FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Tensor Core Units
Haozhi Han, Kun Li, Wei Cui, Donglin Bai, Yiwei Zhang, Liang Yuan, Yifeng Chen, Yunquan Zhang, Ting Cao, Mao Yang
Abstract
While Tensor Core Units (TCUs) excel in AI tasks, their application to HPC algorithms like stencil computations faces significant challenges due to sparsity, which leads to underutilization and exacerbates memory-bound limitations. This paper introduces FlashFFTStencil1, a memory-efficient stencil computing system designed to bridge FFT to fully-dense stencil computations on TCUs. Aimed at bound shifting, FlashFFTStencil comprises three key techniques: Kernel Tailoring on HBM fuses distinct kernels to enhance parallelism while reducing memory transfer and footprint; Architecture Aligning on SMEM restructures FFT-based stencil computations into dense matrix multiplications tailored for shared memory architecture; Computation Streamlining on TCU optimizes TCU utilization and thread parallelism by minimizing pipeline stalls and maximizing register reuse. Notably, a distinctive extension is FlashFFTStencil's ability to enable theoretically unrestricted temporal fusion by FFT. Results show that FlashFFTStencil achieves effective sparsity-free bound shifting, with an average speedup of 2.57x over the state-of-the-art. FlashFFTStencil pioneers a new era in unifying computational patterns within the HPC landscape and bridges them with cutting-edge AI-driven hardware innovations like TCUs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dd6b94d-cfff-44e3-ade4-89bf12279c71Cited by top-tier papers4
- TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPUShixun Wu, Yujia Zhai, Huangliang Dai, Yue Zhu et al.SC 2025 · 5 citations
- SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity TransformationQi Li, Kun Li, Haozhi Han, Liang Yuan et al.SC 2025 · 3 citations
- Characterizing Matrix Multiplication Units across General Parallel Patterns in Scientific ComputingYuechen Lu, Hongwei Zeng, Marc Casas, Weifeng LiuPPoPP 2026 · 1 citation
- SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided SwappingQiqi Gu, Chenpeng Wu, Heng Shi, Jianguo YaoPPoPP 2026 · 1 citation
Builds on7
- FlashFFTConv: Efficient Convolutions for Long Sequences with Tensor CoresDaniel Y. Fu, Hermann Kumbong, Eric Nguyen, Christopher RéICLR 2024 · 41 citations
- ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor CoresYuetao Chen, Kun Li, Yuhao Wang, Donglin Bai et al.PPoPP 2024 · 25 citations
- Improving communication by optimizing on-node data movement with data layoutTuowen Zhao, Mary W. Hall, Hans Johansen, Samuel WilliamsPPoPP 2021 · 19 citations
- Productive Performance Engineering for Weather and Climate Modeling with PythonTal Ben-Nun, Linus Groner, Florian Deconinck, Tobias Wicky et al.SC 2022 · 17 citations
- Scalable Distributed High-Order Stencil ComputationsMathias Jacquelin, Mauricio Araya-Polo, Jie MengSC 2022 · 17 citations
Related papers
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu et al.PPoPP 2025 · 18 citations
- LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor CoresYiwei Zhang, Kun Li, Liang Yuan, Jiawen Cheng et al.SC 2024 · 13 citations
- HStencil: Matrix-Vector Stencil Computation with Interleaved Outer Product and MLAHan Huang, Jiabin Xie, Guangnan Feng, Xianwei Zhang et al.SC 2025 · 5 citations
- FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core ConnectionZiyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo et al.HPCA 2026
- Exploiting Computation Reuse for Stencil AcceleratorsYuze Chi, Jason CongDAC 2020 · 11 citations
