Itoyori: Reconciling Global Address Space and Global Fork-Join Task Parallelism
Shumpei Shiina, Kenjiro Taura
摘要
This paper introduces Itoyori, a task-parallel runtime system designed to tackle the challenge of scaling task parallelism (more specifically, nested fork-join parallelism) beyond a single node. The partitioned global address space (PGAS) model is often employed in task-parallel systems, but naively combining them can lead to poor performance due to fine-grained and redundant remote memory accesses. Itoyori addresses this issue by automatically caching global memory accesses at runtime, enabling efficient cache sharing among parallel tasks running on the same processor. As a real-world case study, we ported an existing task-parallel implementation of the Fast Multipole Method (FMM) to distributed memory with Itoyori and achieved a 7.5× speedup when scaled from a single node to 12 nodes and up to 6.0× faster performance than without caching. This study demonstrates that global-view fork-join programming can be made practical and scalable, while requiring minimal changes to the shared-memory code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Pure: Evolving Message Passing To Better Leverage Shared Memory Within NodesJames Psota, Armando Solar-LezamaPPoPP 2024 · 被引用 1 次
- Enhance the Strong Scaling of LAMMPS on FugakuJianxiong Li, Tong Zhao, Zhuoqiang Guo, Shunchen Shi 等SC 2023 · 被引用 3 次
- Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov methodYasuhiro Idomura, Takuya Ina, Yussuf Ali, Toshiyuki ImamuraSC 2020 · 被引用 6 次
- Sharry: An Efficient and Sharing Far Memory SystemChen Chen, Yuhang Huang, Shuiguang Deng, Jianwei Yin 等DAC 2024
- Pencil: a pipelined algorithm for distributed stencilsHengjie Wang, Aparna ChandramowlishwaranSC 2020 · 被引用 12 次
