Remote Atomic Extension (RAE) for Scalable High Performance Computing
Xi Wang, Brody Williams, John D. Leidel, Alan Ehret, Michel A. Kinsy, Yong Chen
摘要
Emerging data-intensive applications such as graph analytics, machine learning, and data-driven scientific computing are driving the evolution of high-performance computing (HPC) systems from monolithic to scaled-out, heterogeneous, and complex architectures. In these systems, enormous data sets are mapped to discrete nodes to improve the performance of the system by using distributed storage and computing resources. As such, these data distributions induce frequent cross-node data transactions which challenge the performance of large-scale systems. Global atomic operations are one emerging class of the remote data operations that enable lock-free remote shared data operations. However, the cross-node read-modify-write operations consist of multiple distinct data operations and specific atomicity management, which induces a large amount of overhead. As such, these global atomic operations require an efficient communication methodology Existing advanced compo-nents, such as network interface controllers, network fabrics, network-on-chip (NoC) interconnects, are architected together to improve the system performance. However, complex software infrastructures are needed to provide integration between each discrete component. As a result, the redundant software routines across distinct devices induce a large amount of overhead that causes performance degradationIn this paper, we propose a remote atomic extension (RAE) design that provides inherent ISA-level instructions and micro-architecture support for remote atomic operations based on the RISC-V instruction set architecture (ISA). We design a toolchain and evaluate the RAE infrastructure via simulation. Our experiment results show that RAE eliminates 89.71% of the redundant software instructions used for remote atomic accesses and improves the performance by 17.61% on average (up to 23.35%), compared with the OpenSHMEM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic OperationsYicong Zhang, Mingyu Wang, Wangguang Wang, Yangzhan Mai 等MICRO 2024 · 被引用 4 次
- Only Buffer When You Need To: Reducing On-chip GPU Traffic with Reconfigurable Local Atomic BuffersPreyesh Dalmia, Rohan Mahapatra, Matthew D. SinclairHPCA 2022 · 被引用 11 次
- ATUNs: Modular and Scalable Support for Atomic Operations in a Shared Memory MultiprocessorAndreas Kurth, Samuel Riedel, Florian Zaruba, Torsten Hoefler 等DAC 2020 · 被引用 5 次
- ARCANE: Adaptive RISC-V Cache Architecture for Near-memory ExtensionsVincenzo Petrolo, Flavia Guella, Michele Caon, Pasquale Davide Schiavone 等DAC 2025 · 被引用 1 次
- Breaking Barriers in Atomic Scaling: A Hardware-Software-Collaborated Framework to Deconstruct RDMA AtomicGuangyang Deng, Qiangsheng Su, Zhirong Shen, Qing Wang 等ISCA 2026
