SC2020Top-tier venue
Pencil: a pipelined algorithm for distributed stencils
Hengjie Wang, Aparna Chandramowlishwaran
Abstract
Stencil computations are at the core of various Computational Fluid Dynamics (CFD) applications and have been well-studied for several decades. Typically they're highly memory-bound and as a result, numerous tiling algorithms have been proposed to improve its performance. Although efficient, most of these algorithms are designed for single iteration spaces on shared-memory machines. However, in CFD, we are confronted with multi-block structured girds composed of multiple connected iteration spaces distributed across many nodes.In this paper, we propose a pipelined stencil algorithm called Pencil for distributed memory machines that applies to practical CFD problems that span multiple iteration spaces. Based on an in-depth analysis of cache tiling on a single node, we first identify both the optimal combination of MPI and OpenMP for temporal tiling and the best tiling approach, which outperforms the state-of-the-art automatic parallelization tool Pluto by up to . Then, we adopt DeepHalo to decouple the multiple connected iteration spaces so that temporal tiling can be applied to each space. Finally, we achieve overlap by pipelining the computation and communication without sacrificing the advantage from temporal cache tiling. Pencil is evaluated using 4 stencils across 6 numerical schemes on two distributed memory machines with Omni-Path and InfiniBand networks. On the Omni-Path system, Pencil exhibits outstanding weak and strong scalability for up to 128 nodes and outperforms MPI+OpenMP Funneled with space tiling by on a multi-block grid with 32 nodes.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 73a4f395-be15-4310-b2a0-c3b69b248cecCited by top-tier papers4
- Breaking the computation and communication abstraction barrier in distributed machine learning workloadsAbhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet et al.ASPLOS 2022 · 68 citations
- A shared compilation stack for distributed-memory parallelism in stencil DSLsGeorge Bisbas, Anton Lydike, Emilien Bauer, Nick Brown et al.ASPLOS 2024 · 11 citations
- Lessons Learned on MPI+Threads CommunicationRohit Zambre, Aparna ChandramowlishwaranSC 2022 · 5 citations
- Breaking Boundaries: Distributed Domain Decomposition with Scalable Physics-Informed Neural PDE SolversArthur Feeney, Zitong Li, Ramin Bostanabad, Aparna ChandramowlishwaranSC 2023 · 2 citations
Related papers
- Towards Scalable Unstructured Mesh Computations on Shared Memory Many-CoresHaozhong Qiu, Chuanfu Xu, Jianbin Fang, Liang Deng et al.PPoPP 2024 · 8 citations
- Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured GridHuanqi Cao, Shizhi Tang, Qianchao Zhu, Bowen Yu et al.OOPSLA 2023 · 4 citations
- Reducing redundancy in data organization and arithmetic calculation for stencil computationsKun Li, Liang Yuan, Yunquan Zhang, Yue YueSC 2021 · 12 citations
- Telos: A Dataflow Accelerator for Sparse Triangular Solver of Partial Differential EquationsXiaochen Hao, Hao Luo, Chu Wang, Chao Yang et al.ISCA 2025
- Itoyori: Reconciling Global Address Space and Global Fork-Join Task ParallelismShumpei Shiina, Kenjiro TauraSC 2023 · 6 citations
