Optimizing the Bruck Algorithm for Non-uniform All-to-all Communication
Ke Fan, Thomas Gilray, Valerio Pascucci, Xuan Huang, Kristopher K. Micinski, Sidharth Kumar
摘要
In MPI, collective routines MPI_Alltoall and MPI_Alltoallv play an important role in facilitating all-to-all inter-process data exchange. MPI_Alltoallv is a generalization of MPI_Alltoall, supporting the exchange of non-uniform distributions of data. Popular implementations of MPI, such as MPICH and OpenMPI, implement MPI_Alltoall using a combination of techniques such as the Spread-out algorithm and the Bruck algorithm. Spread-out has a linear complexity in P, compared to Bruck's logarithmic complexity (P: process count); a selection between these two techniques is made at runtime based on the data block size. However, MPI_Alltoallv is typically implemented using only variants of the spread-out algorithm, and therefore misses out on the performance benefits that the log-time Bruck algorithm offers (especially for smaller data loads).
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node CommunicationsXi Wang, Bin Ma, Jongryool Kim, Byungil Koh 等SC 2025 · 被引用 5 次
- Datalog with First-Class FactsThomas Gilray, Arash Sahebolamri, Yihao Sun, Sowmith Kunapaneni 等VLDB 2025 · 被引用 3 次
- Optimizing Datalog for the GPUYihao Sun, Ahmedur Rahman Shovon, Thomas Gilray, Sidharth Kumar 等ASPLOS 2025 · 被引用 3 次
- Parameterized Algorithms for Non-uniform All-to-allKe Fan, Jens Domke, Seydou Ba, Sidharth KumarHPDC 2025 · 被引用 2 次
相关 Paper
- Efficient all-to-all Collective Communication Schedules for Direct-connect TopologiesPrithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal 等HPDC 2024 · 被引用 7 次
- Optimizing MPI Collectives on Shared Memory Multi-CoresJintao Peng, Jianbin Fang, Jie Liu, Min Xie 等SC 2023 · 被引用 10 次
- Improving all-to-many personalized communication in two-phase I/OQiao Kang, Robert B. Ross, Robert Latham, Sunwoo Lee 等SC 2020 · 被引用 12 次
- FAST: An Efficient Scheduler for All-to-All GPU CommunicationYiran Lei, Dongjoo Lee, Liangyu Zhao, Daniar Kurniawan 等NSDI 2026 · 被引用 13 次
- Distributed quantum computing with QMPIThomas Häner, Damian S. Steiger, Torsten Hoefler, Matthias TroyerSC 2021 · 被引用 48 次
