ODOS-MPI: HPC-Friendly SmartNIC Offloading of Computation/Communication Kernels
Muhammad Usman, Mariano Benito, Sergio Iserte, Antonio J. Peña
摘要
The increasing complexity and scale of high-performance computing (HPC) workloads demand innovative approaches to optimize both computation and communication. While OpenMP has been widely adopted for intra-node parallelism and MPI for inter-node communication, emerging SmartNICs introduce new opportunities for offloading communication-intensive tasks. In this work, we extend OpenMP to support MPI kernel offloading to SmartNICs. Our implementation integrates Open MPI communication offloading into the LLVM compiler while utilizing DOCA SDK for efficient interaction with Nvidia BlueField DPUs. Leveraging OpenMP eliminates the need for direct low-level programming, lowering the entry barrier for domain scientists. We demonstrate our framework’s versatility by implementing a SmartNIC-enabled version of the MPI OSU micro-benchmarks and improving the execution time of an atmospheric weather simulation by over 18%, thanks to concurrent computation and communication.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Characterizing Off-path SmartNIC for Accelerating Distributed SystemsXingda Wei, Rongxin Cheng, Yuhan Yang, Rong Chen 等OSDI 2023 · 被引用 68 次
- Conspirator: SmartNIC-Aided Control Plane for Distributed ML WorkloadsYunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao 等USENIX ATC 2024 · 被引用 13 次
- Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AIMikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek, Rami Nudelman 等SC 2024 · 被引用 15 次
- NutCracker: A Compilation Framework for Hybrid DPU ArchitecturesYihan Yang, Haifeng Sun, Antoine Kaufmann, Jialin LiEuroSys 2026 · 被引用 2 次
- Non-recurring engineering (NRE) best practices: a case study with the NERSC/NVIDIA OpenMP contractChristopher S. Daley, Annemarie Southwell, Rahulkumar Gayatri, Scott Biersdorfff 等SC 2021 · 被引用 2 次
