Improving communication by optimizing on-node data movement with data layout
Tuowen Zhao, Mary W. Hall, Hans Johansen, Samuel Williams
Abstract
We present optimizations to improve communication performance by reducing on-node data movement for a class of distributed memory applications. The primary concept is to eliminate the data movement associated with packing and unpacking subsets of the data during communication. With the rapid rise in network injection bandwidth reducing off-node data movement cost, on-node data movement can be significantly more expensive than computation and network communication. This data movement is especially costly for small domains - as in memory-intensive multi-physics codes or when strong scaling to reduce time-to-solution. The optimizations presented include (1) optimizing data layout through indirection to enable pack-free communication; (2) creating contiguous views of memory using memory mapping thus minimizing the number of messages; and (3) applying these techniques to intra-node data movement including CPU-GPU data movement. The benefits of these optimizations are demonstrated in stencil benchmarks against a highly-optimized baseline, reducing communication time by up to 14.4×.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1aa55131-dd77-450a-82da-0262df04d44cCited by top-tier papers6
- ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor CoresYuetao Chen, Kun Li, Yuhao Wang, Donglin Bai et al.PPoPP 2024 · 25 citations
- LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor CoresYiwei Zhang, Kun Li, Liang Yuan, Jiawen Cheng et al.SC 2024 · 13 citations
- A shared compilation stack for distributed-memory parallelism in stencil DSLsGeorge Bisbas, Anton Lydike, Emilien Bauer, Nick Brown et al.ASPLOS 2024 · 11 citations
- FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Tensor Core UnitsHaozhi Han, Kun Li, Wei Cui, Donglin Bai et al.PPoPP 2025 · 7 citations
- SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity TransformationQi Li, Kun Li, Haozhi Han, Liang Yuan et al.SC 2025 · 3 citations
Related papers
- CAB-MPI: exploring interprocess work-stealing towards balanced MPI communicationKaiming Ouyang, Min Si, Atsushi Hori, Zizhong Chen et al.SC 2020 · 11 citations
- Pure: Evolving Message Passing To Better Leverage Shared Memory Within NodesJames Psota, Armando Solar-LezamaPPoPP 2024 · 1 citation
- MGI: A Communication Framework for Data Processing in Massive GPU InfrastructuresDi Wu, Hongshi Tan, Hanzhang Yang, Bingsheng He et al.VLDB 2026
- The Memory Processing Unit: A Generalized Interface for End-to-End In-Memory ExecutionMinh S. Q. Truong, Yiqiu Sun, Dawei Xiong, Amol Shah et al.HPCA 2026 · 1 citation
- Dynamic Detection of Inefficient Data Mapping Patterns in Heterogeneous OpenMP ApplicationsLuke Marzen, Junhyung Shim, Ali JannesariPPoPP 2026
