Evaluation of a minimally synchronous algorithm for 2: 1 octree balance
Hansol Suh, Tobin Isaac
摘要
The p4est library implements octree-based adaptive mesh refinement (AMR) and has demonstrated parallel scalability beyond 100,000 MPI processes in previous weak scaling studies. This work focuses on the strong scalability of mesh adaptivity in p4est, where the communication pattern of the existing 2:1-balance is a latency bottleneck. The sorting-based algorithm of Malhotra and Biros has balanced communication, but synchronizes all processes. We propose an algorithm that combines sorting and neighbor-to-neighbor exchange to minimize the number of processes each process synchronizes with. We measure the performance of these algorithms on several test problems on Stampede2 at TACC. Both the parallel-sorting and minimally-synchronous algorithms significantly outperform the existing algorithm and have nearly identical performance out to 1,024 Xeon Phi KNL nodes, meaning the asymptotic advantage of the minimally-synchronous algorithm does not translate to improved performance at this scale. We conclude by showing that global metadata communication will limit future strong scaling.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Enhance the Strong Scaling of LAMMPS on FugakuJianxiong Li, Tong Zhao, Zhuoqiang Guo, Shunchen Shi 等SC 2023 · 被引用 3 次
- AMRaCut: Scalable Partitioning for Adaptive Mesh RefinementBudvin Edippuliarachchi, David Van Komen, Hari SundarSC 2025 · 被引用 1 次
- Improving all-to-many personalized communication in two-phase I/OQiao Kang, Robert B. Ross, Robert Latham, Sunwoo Lee 等SC 2020 · 被引用 12 次
- Towards Scalable Unstructured Mesh Computations on Shared Memory Many-CoresHaozhong Qiu, Chuanfu Xu, Jianbin Fang, Liang Deng 等PPoPP 2024 · 被引用 8 次
- Pure: Evolving Message Passing To Better Leverage Shared Memory Within NodesJames Psota, Armando Solar-LezamaPPoPP 2024 · 被引用 1 次
