Pencil: a pipelined algorithm for distributed stencils
Hengjie Wang, Aparna Chandramowlishwaran
摘要
Stencil computations are at the core of various Computational Fluid Dynamics (CFD) applications and have been well-studied for several decades. Typically they're highly memory-bound and as a result, numerous tiling algorithms have been proposed to improve its performance. Although efficient, most of these algorithms are designed for single iteration spaces on shared-memory machines. However, in CFD, we are confronted with multi-block structured girds composed of multiple connected iteration spaces distributed across many nodes.In this paper, we propose a pipelined stencil algorithm called Pencil for distributed memory machines that applies to practical CFD problems that span multiple iteration spaces. Based on an in-depth analysis of cache tiling on a single node, we first identify both the optimal combination of MPI and OpenMP for temporal tiling and the best tiling approach, which outperforms the state-of-the-art automatic parallelization tool Pluto by up to . Then, we adopt DeepHalo to decouple the multiple connected iteration spaces so that temporal tiling can be applied to each space. Finally, we achieve overlap by pipelining the computation and communication without sacrificing the advantage from temporal cache tiling. Pencil is evaluated using 4 stencils across 6 numerical schemes on two distributed memory machines with Omni-Path and InfiniBand networks. On the Omni-Path system, Pencil exhibits outstanding weak and strong scalability for up to 128 nodes and outperforms MPI+OpenMP Funneled with space tiling by on a multi-block grid with 32 nodes.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Breaking the computation and communication abstraction barrier in distributed machine learning workloadsAbhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet 等ASPLOS 2022 · 被引用 68 次
- A shared compilation stack for distributed-memory parallelism in stencil DSLsGeorge Bisbas, Anton Lydike, Emilien Bauer, Nick Brown 等ASPLOS 2024 · 被引用 11 次
- Lessons Learned on MPI+Threads CommunicationRohit Zambre, Aparna ChandramowlishwaranSC 2022 · 被引用 5 次
- Breaking Boundaries: Distributed Domain Decomposition with Scalable Physics-Informed Neural PDE SolversArthur Feeney, Zitong Li, Ramin Bostanabad, Aparna ChandramowlishwaranSC 2023 · 被引用 2 次
相关 Paper
- Towards Scalable Unstructured Mesh Computations on Shared Memory Many-CoresHaozhong Qiu, Chuanfu Xu, Jianbin Fang, Liang Deng 等PPoPP 2024 · 被引用 8 次
- Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured GridHuanqi Cao, Shizhi Tang, Qianchao Zhu, Bowen Yu 等OOPSLA 2023 · 被引用 4 次
- Reducing redundancy in data organization and arithmetic calculation for stencil computationsKun Li, Liang Yuan, Yunquan Zhang, Yue YueSC 2021 · 被引用 12 次
- Telos: A Dataflow Accelerator for Sparse Triangular Solver of Partial Differential EquationsXiaochen Hao, Hao Luo, Chu Wang, Chao Yang 等ISCA 2025
- Itoyori: Reconciling Global Address Space and Global Fork-Join Task ParallelismShumpei Shiina, Kenjiro TauraSC 2023 · 被引用 6 次
