Lessons Learned on MPI+Threads Communication
Rohit Zambre, Aparna Chandramowlishwaran
摘要
Hybrid MPI+threads programming is gaining prominence, but, in practice, applications perform slower with it compared to the MPI everywhere model. The most critical challenge to the parallel efficiency of MPI+threads applications is slow MPI_THREAD_MULTIPLE performance. MPI libraries have recently made significant strides on this front, but to exploit their capabilities, users must expose the communication parallelism in their MPI+threads applications. Recent studies show that MPI 4.0 provides users with new performance-oriented options to do so, but our evaluation of these new mechanisms shows that they pose several challenges. An alternative design is MPI Endpoints. In this paper, we present a comparison of the different designs from the perspective of MPI's end-users: domain scientists and application developers. We evaluate the mechanisms on metrics beyond performance such as usability, scope, and portability. Based on the lessons learned, we make a case for a future direction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Improving all-to-many personalized communication in two-phase I/OQiao Kang, Robert B. Ross, Robert Latham, Sunwoo Lee 等SC 2020 · 被引用 12 次
- KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPITim Niklas Uhl, Matthias Schimek, Lukas Hübner, Demian Hespe 等SC 2024 · 被引用 12 次
- Graphite: A NUMA-aware HPC System for Graph Analytics Based on a new MPI * X Parallelism ModelMohammad Hasanzadeh-Mofrad, Rami G. Melhem, Muhammad Yousuf Ahmad, Mohammad HammoudVLDB 2020 · 被引用 142 次
- Embracing Irregular Parallelism in HPC with YGMTrevor Steil, Tahsin Reza, Benjamin Priest, Roger PearceSC 2023 · 被引用 8 次
- swKokkos: An Athread Backend for Enhanced Kokkos with the Sunway Heterogeneous ArchitectureJunlin Wei, Jinrong Jiang, Wu Wang, Chen Li 等EuroSys 2026
