Improving communication by optimizing on-node data movement with data layout
Tuowen Zhao, Mary W. Hall, Hans Johansen, Samuel Williams
摘要
We present optimizations to improve communication performance by reducing on-node data movement for a class of distributed memory applications. The primary concept is to eliminate the data movement associated with packing and unpacking subsets of the data during communication. With the rapid rise in network injection bandwidth reducing off-node data movement cost, on-node data movement can be significantly more expensive than computation and network communication. This data movement is especially costly for small domains - as in memory-intensive multi-physics codes or when strong scaling to reduce time-to-solution. The optimizations presented include (1) optimizing data layout through indirection to enable pack-free communication; (2) creating contiguous views of memory using memory mapping thus minimizing the number of messages; and (3) applying these techniques to intra-node data movement including CPU-GPU data movement. The benefits of these optimizations are demonstrated in stencil benchmarks against a highly-optimized baseline, reducing communication time by up to 14.4×.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor CoresYuetao Chen, Kun Li, Yuhao Wang, Donglin Bai 等PPoPP 2024 · 被引用 25 次
- LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor CoresYiwei Zhang, Kun Li, Liang Yuan, Jiawen Cheng 等SC 2024 · 被引用 13 次
- A shared compilation stack for distributed-memory parallelism in stencil DSLsGeorge Bisbas, Anton Lydike, Emilien Bauer, Nick Brown 等ASPLOS 2024 · 被引用 11 次
- FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Tensor Core UnitsHaozhi Han, Kun Li, Wei Cui, Donglin Bai 等PPoPP 2025 · 被引用 7 次
- SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity TransformationQi Li, Kun Li, Haozhi Han, Liang Yuan 等SC 2025 · 被引用 3 次
相关 Paper
- CAB-MPI: exploring interprocess work-stealing towards balanced MPI communicationKaiming Ouyang, Min Si, Atsushi Hori, Zizhong Chen 等SC 2020 · 被引用 11 次
- Pure: Evolving Message Passing To Better Leverage Shared Memory Within NodesJames Psota, Armando Solar-LezamaPPoPP 2024 · 被引用 1 次
- MGI: A Communication Framework for Data Processing in Massive GPU InfrastructuresDi Wu, Hongshi Tan, Hanzhang Yang, Bingsheng He 等VLDB 2026
- The Memory Processing Unit: A Generalized Interface for End-to-End In-Memory ExecutionMinh S. Q. Truong, Yiqiu Sun, Dawei Xiong, Amol Shah 等HPCA 2026 · 被引用 1 次
- Dynamic Detection of Inefficient Data Mapping Patterns in Heterogeneous OpenMP ApplicationsLuke Marzen, Junhyung Shim, Ali JannesariPPoPP 2026
