Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured Grid
Huanqi Cao, Shizhi Tang, Qianchao Zhu, Bowen Yu, Wenguang Chen
摘要
Partial differential equation (PDE) solvers are extensively utilized across numerous scientific and engineering fields. However, achieving high performance and scalability often necessitates intricate and low-level programming, particularly when leveraging deterministic sparsity patterns in structured grids.
In this paper, we propose an innovative domain-specific language (DSL), Mat2Stencil, with its compiler, for PDE solvers on structured grids. Mat2Stencil introduces a structured sparse matrix abstraction, facilitating modular, flexible, and easy-to-use expression of solvers across a broad spectrum, encompassing components such as Jacobi or Gauss-Seidel preconditioners, incomplete LU or Cholesky decompositions, and multigrid methods built upon them. Our DSL compiler subsequently generates matrix-free code consisting of generalized stencils through multi-stage programming. The code allows spatial loop-carried dependence in the form of quasi-affine loops, in addition to the Jacobi-style stencil's embarrassingly parallel on spatial dimensions. We further propose a novel automatic parallelization technique for the spatially dependent loops, which offers a compile-time deterministic task partitioning for threading, calculates necessary inter-thread synchronization automatically, and generates an efficient multi-threaded implementation with fine-grained synchronization.
Implementing 4 benchmarking programs, 3 of them being the pseudo-applications in NAS Parallel Benchmarks with 6.3% lines of code and 1 being matrix-free High Performance Conjugate Gradients with 16.4% lines of code, we achieve up to 1.67× and on average 1.03× performance compared to manual implementations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- Enabling and scaling the HPCG benchmark on the newest generation Sunway supercomputer with 42 million heterogeneous coresQianchao Zhu, Hao Luo, Chao Yang, Mingshuo Ding 等SC 2021 · 被引用 44 次
- FreeTensor: a free-form DSL with holistic optimizations for irregular tensor programsShizhi Tang, Jidong Zhai, Haojie Wang, Lin Jiang 等PLDI 2022 · 被引用 16 次
相关 Paper
- Telos: A Dataflow Accelerator for Sparse Triangular Solver of Partial Differential EquationsXiaochen Hao, Hao Luo, Chu Wang, Chao Yang 等ISCA 2025
- A shared compilation stack for distributed-memory parallelism in stencil DSLsGeorge Bisbas, Anton Lydike, Emilien Bauer, Nick Brown 等ASPLOS 2024 · 被引用 11 次
- Pencil: a pipelined algorithm for distributed stencilsHengjie Wang, Aparna ChandramowlishwaranSC 2020 · 被引用 12 次
- Stencil-Lifting: Hierarchical Recursive Lifting System for Extracting Summary of Stencil Kernel in Legacy CodesMingyi Li, Junmin Xiao, Siyan Chen, Hui Ma 等OOPSLA 2025
- MeshFEM: A Block-accelerated Solver for Nonlinear Finite ElementsHaleh Mohammadian, Xinzhuo Hu, Roi Poranne, Julian PanettaSIGGRAPH 2026 · 被引用 1 次
