Embracing Irregular Parallelism in HPC with YGM
Trevor Steil, Tahsin Reza, Benjamin Priest, Roger Pearce
摘要
YGM is a general-purpose asynchronous distributed computing library for C++/MPI, designed to handle the irregular data access patterns and small messages of graph algorithms and data science applications. It uses data serialization to give an easily usable active message interface and message aggregation to maximize application throughput. Our design philosophy makes a tradeoff that increases network bandwidth utilization at the cost of added latency. We provide a suite of benchmarks showcasing YGM's performance. Compared to similar distributed active message benchmark implementations that do not provide message buffering, we are able to achieve over 10x throughput on thousands of cores at a latency cost that can be as small as 2x or as large as 100x, depending on the machine being used. For applications that can be written to be latency-tolerant, this represents a significant potential performance improvement through using YGM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- AGILE: Lightweight and Efficient Asynchronous GPU-SSD IntegrationZhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones 等SC 2025 · 被引用 3 次
- MGI: A Communication Framework for Data Processing in Massive GPU InfrastructuresDi Wu, Hongshi Tan, Hanzhang Yang, Bingsheng He 等VLDB 2026
- COSMOS: Performance Portable Graph Pattern Matching with Domain-Specific Software Distributed Shared MemoryZhiheng Lin, Ke Meng, Changjie Xu, Weichen Cao 等SC 2025 · 被引用 1 次
- KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPITim Niklas Uhl, Matthias Schimek, Lukas Hübner, Demian Hespe 等SC 2024 · 被引用 12 次
- GraphCom: Communication Hierarchy-aware Graph Engine for Distributed Model TrainingXinbiao Gan, Tiejun Li, Liang Wu, Qiang Zhang 等WWW 2025 · 被引用 1 次
