Lune

SC2024Top-tier venue

Exploring Efficient Partial Differential Equation Solution Using Speed Galerkin Transformer

Xun Wang, Zeyang Zhu, Xiangyu Meng, Tao Song

2024Year

Abstract

Fourier Neural Operator (FNO) has been proven to be a universal and effective deep learning framework capable of achieving remarkable accuracy on Partial Differential Equation (PDE) solution problem. However, certain key components of emerging FNO-based models cannot leverage hardware potential, which makes it difficult to apply in high resolution and high realtime demand scenario. This paper presents a high optimized model called Speed Galerkin Transformer, including multilevel parallel SliceK-SplitK-ReduceK strategy for batched skinny matrix multiplication, memory layout optimization for QKV matrices and positional encodings and multi-head layer normalization fusion, as well as batched transposition optimization with strided scattering and gathering in 2D FNO, and these strategies can achieve 10.29x,4.41x10.29 \mathrm{x}, 4.41 \mathrm{x} and 2.38 x speedup respectively under specific configuration. When solving the Darcy Flow equation at 512x512 resolution, the Speed Galerkin Transformer model can achieve about 1.72 x speedup, and achieve more than 90%\mathbf{9 0 \%} parallel efficiency on 8 GPUs.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 5e725d58-0eea-4959-b6dc-551df7e0e82f

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines