SC2024Top-tier venue
LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor Cores
Yiwei Zhang, Kun Li, Liang Yuan, Jiawen Cheng, Yunquan Zhang, Ting Cao, Mao Yang
Abstract
Stencil computations play a pivotal role in numerous scientific and industrial applications, yet their efficient execution on specialized hardware accelerators like Tensor Core Units (TCUs) remains a challenge. Despite previous attempts to address this issue, performance bottlenecks persist, particularly in memory access redundancy. This paper introduces LoRAStencil 1 , a novel stencil computing system designed to mitigate memory access redundancies on TCUs through low-rank adaptation. We first identify a nuanced form of this redundancy, dimension residue, specific to TCUs. Then LoRAStencil leverages orchestrated mathematical transformations to decompose stencil weight matrices into smaller rank-1 matrices, facilitating efficient data gathering along residual dimensions. It comprises three key components: memory-efficient Residual Dimension Gathering to facilitate more data reuse, compute-saving Pyramidal Matrix Adaptation to exploit the inherent low-rank characteristics, and performance-boosting Butterfly Vector Swapping to circumvent all data shuffles. Comprehensive evaluations demonstrate that LoRAStencil address dimension residues effectively, which outperforms state-of-the-arts with up to a 2.16x speedup, offering promising advancements for efficient tensorized stencil computation on TCUs by Low-Rank Adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bad7844e-a21d-4d3a-b811-62fd9765cd0aCited by top-tier papers7
- FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Tensor Core UnitsHaozhi Han, Kun Li, Wei Cui, Donglin Bai et al.PPoPP 2025 · 7 citations
- HStencil: Matrix-Vector Stencil Computation with Interleaved Outer Product and MLAHan Huang, Jiabin Xie, Guangnan Feng, Xianwei Zhang et al.SC 2025 · 5 citations
- SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity TransformationQi Li, Kun Li, Haozhi Han, Liang Yuan et al.SC 2025 · 3 citations
- Jigsaw: Toward Conflict-free Vectorized Stencil Computation by Tessellating Swizzled RegistersYiwei Zhang, Kun Li, Liang Yuan, Haozhi Han et al.PPoPP 2025 · 2 citations
- Characterizing Matrix Multiplication Units across General Parallel Patterns in Scientific ComputingYuechen Lu, Hongwei Zeng, Marc Casas, Weifeng LiuPPoPP 2026 · 1 citation
Builds on7
- AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstractionSize Zheng, Renze Chen, Anjiang Wei, Yicheng Jin et al.ISCA 2022 · 63 citations
- ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor CoresYuetao Chen, Kun Li, Yuhao Wang, Donglin Bai et al.PPoPP 2024 · 25 citations
- Improving communication by optimizing on-node data movement with data layoutTuowen Zhao, Mary W. Hall, Hans Johansen, Samuel WilliamsPPoPP 2021 · 19 citations
- Productive Performance Engineering for Weather and Climate Modeling with PythonTal Ben-Nun, Linus Groner, Florian Deconinck, Tobias Wicky et al.SC 2022 · 17 citations
- Scalable Distributed High-Order Stencil ComputationsMathias Jacquelin, Mauricio Araya-Polo, Jie MengSC 2022 · 17 citations
Related papers
- SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided SwappingQiqi Gu, Chenpeng Wu, Heng Shi, Jianguo YaoPPoPP 2026 · 1 citation
- Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor CoresHyeonjin Kim, Sungwoo Ahn, Yunho Oh, Bogil Kim et al.MICRO 2020 · 27 citations
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu et al.PPoPP 2025 · 18 citations
- LoRAFusion: Efficient LoRA Fine-Tuning for LLMsZhanda Zhu, Qidong Su, Yaoyao Ding, Kevin Song et al.EuroSys 2026 · 2 citations
- Temporal vectorization for stencilsLiang Yuan, Hang Cao, Yunquan Zhang, Kun Li et al.SC 2021 · 11 citations
