SC2020Top-tier venue
CAB-MPI: exploring interprocess work-stealing towards balanced MPI communication
Kaiming Ouyang, Min Si, Atsushi Hori, Zizhong Chen, Pavan Balaji
Abstract
Load balance is essential for high-performance applications. Unbalanced communication can cause severe performance degradation, even in computation-balanced BSP applications. Designing communication-balanced applications is challenging, however, because of the diverse communication implementations at the underlying runtime system. In this paper, we address this challenge through an interprocess workstealing scheme based on process-memory-sharing techniques. We present CAB-MPI, an MPI implementation that can identify idle processes inside MPI and use these idle resources to dynamically balance communication workload on the node. We design throughput-optimized strategies to ensure efficient stealing of the data movement tasks. We demonstrate the benefit of work stealing through several internal processes in MPI, including intranode data transfer, pack/unpack for noncontiguous communication, and computation in one-sided accumulates. The implementation is evaluated through a set of microbenchmarks and proxy applications on Intel Xeon and Xeon Phi platforms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e9ddce9-e105-49d6-a86a-a2cf02e0c424Cited by top-tier papers1
Ask how each one uses itRelated papers
- Improving communication by optimizing on-node data movement with data layoutTuowen Zhao, Mary W. Hall, Hans Johansen, Samuel WilliamsPPoPP 2021 · 19 citations
- Pure: Evolving Message Passing To Better Leverage Shared Memory Within NodesJames Psota, Armando Solar-LezamaPPoPP 2024 · 1 citation
- cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node CommunicationsXi Wang, Bin Ma, Jongryool Kim, Byungil Koh et al.SC 2025 · 5 citations
- Merchandiser: Data Placement on Heterogeneous Memory for Task-Parallel HPC Applications with Load-Balance AwarenessZhen Xie, Jie Liu, Jiajia Li, Dong LiPPoPP 2023 · 18 citations
- A Programming Model for GPU Load BalancingMuhammad Osama, Serban D. Porumbescu, John D. OwensPPoPP 2023 · 10 citations
