A Scalable Hybrid Total FETI Method for Massively Parallel FEM Simulations
Kehao Lin, Chunbao Zhou, Yan Zeng, Ningming Nie, Jue Wang, Shigang Li, Yangde Feng, Yangang Wang, Kehan Yao, Tiechui Yao, Jilin Zhang, Jian Wan
摘要
The Hybrid Total Finite Element Tearing and Interconnecting (HTFETI) method plays an important role in solving large-scale and complex engineering problems. This method needs to handle numerous matrix-vector multiplications. Directly calling the vendor-optimized library for general matrix-vector multiplication (gemv) on GPU leads to low performance, since it does not consider optimizations for different matrix sizes in HTFETI, i.e. different row and column sizes. In addition, state-of-the-art graph partitioning methods cannot guarantee load balancing for HTFETI, since the matrix size is determined by the length of the subdomain boundary. To solve the problems above, we first port gemv to the multi-stream pipeline scheme and develop a new batched kernel function on GPU, which brings 15% 30% throughput improvement and 37% average GFLOPs improvement, respectively. We also propose a multi-grained load-balancing scheme based on graph repartitioning and work-stealing, and the load imbalance ratio is down to 1.05 1.09 from 1.5. We have successfully applied the scalable HTFETI method to simulate the whole core assembly of China Experimental Fast Reactor (CEFR) for steady-state analysis, and the efficiencies of weak scalability and strong scalability reach 78% and 72% on 12,288 GPUs, respectively. As far as we know, this is the first time that HTFETI has been used in large-scale and high-fidelity whole core assembly simulation.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Large-Scale Simulation of Structural Dynamics Computing on GPU ClustersYumeng Shi, Ningming Nie, Jue Wang, Kehao Lin 等SC 2023 · 被引用 7 次
- Scaling the hartree-fock matrix build on summitGiuseppe M. J. Barca, David L. Poole, Jorge L. Galvez Vallejo, Melisa Alkan 等SC 2020 · 被引用 29 次
- Full-Core Fluid-Structure-Interaction Simulation of Nuclear Reactor on CPU+GPU Hybrid ClustersXue Miao, Jue Wang, Qida Lin, Shufei Zhang 等HPDC 2026
- Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov methodYasuhiro Idomura, Takuya Ina, Yussuf Ali, Toshiyuki ImamuraSC 2020 · 被引用 6 次
- ANT-MOC: Scalable Neutral Particle Transport Using 3D Method of Characteristics on Multi-GPU SystemsShunde Li, Zongguo Wang, Lingkun Bu, Jue Wang 等SC 2023 · 被引用 7 次
