Optimizing the Bruck Algorithm for Non-uniform All-to-all Communication
Ke Fan, Thomas Gilray, Valerio Pascucci, Xuan Huang, Kristopher K. Micinski, Sidharth Kumar
Abstract
In MPI, collective routines MPI_Alltoall and MPI_Alltoallv play an important role in facilitating all-to-all inter-process data exchange. MPI_Alltoallv is a generalization of MPI_Alltoall, supporting the exchange of non-uniform distributions of data. Popular implementations of MPI, such as MPICH and OpenMPI, implement MPI_Alltoall using a combination of techniques such as the Spread-out algorithm and the Bruck algorithm. Spread-out has a linear complexity in P, compared to Bruck's logarithmic complexity (P: process count); a selection between these two techniques is made at runtime based on the data block size. However, MPI_Alltoallv is typically implemented using only variants of the spread-out algorithm, and therefore misses out on the performance benefits that the log-time Bruck algorithm offers (especially for smaller data loads).
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cacc08a4-203e-4cf5-a019-5a64625658feCited by top-tier papers4
- cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node CommunicationsXi Wang, Bin Ma, Jongryool Kim, Byungil Koh et al.SC 2025 · 5 citations
- Datalog with First-Class FactsThomas Gilray, Arash Sahebolamri, Yihao Sun, Sowmith Kunapaneni et al.VLDB 2025 · 3 citations
- Optimizing Datalog for the GPUYihao Sun, Ahmedur Rahman Shovon, Thomas Gilray, Sidharth Kumar et al.ASPLOS 2025 · 3 citations
- Parameterized Algorithms for Non-uniform All-to-allKe Fan, Jens Domke, Seydou Ba, Sidharth KumarHPDC 2025 · 2 citations
Related papers
- Efficient all-to-all Collective Communication Schedules for Direct-connect TopologiesPrithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal et al.HPDC 2024 · 7 citations
- Optimizing MPI Collectives on Shared Memory Multi-CoresJintao Peng, Jianbin Fang, Jie Liu, Min Xie et al.SC 2023 · 10 citations
- Improving all-to-many personalized communication in two-phase I/OQiao Kang, Robert B. Ross, Robert Latham, Sunwoo Lee et al.SC 2020 · 12 citations
- FAST: An Efficient Scheduler for All-to-All GPU CommunicationYiran Lei, Dongjoo Lee, Liangyu Zhao, Daniar Kurniawan et al.NSDI 2026 · 13 citations
- Distributed quantum computing with QMPIThomas Häner, Damian S. Steiger, Torsten Hoefler, Matthias TroyerSC 2021 · 48 citations
