OrderLight: Lightweight Memory-Ordering Primitive for Efficient Fine-Grained PIM Computations
Anirban Nag, Rajeev Balasubramonian
摘要
Modern workloads such as neural networks, genomic analysis, and data analytics exhibit significant data-intensive phases (low compute to byte ratio) and, as such, stand to gain considerably by using processing-in-memory (PIM) solutions along with more traditional accelerators. While PIM has been researched extensively, the granularity of computation offload to PIM and the granularity of memory access arbitration between host and PIM, as well as their implications, have received relatively little attention. In this work, we first introduce a taxonomy to study the design space whilst considering these two aspects. Based on this taxonomy, we observe that much of PIM research to date has largely relied on coarse-grained approaches which, we argue, have steep costs (incompatibility with mainstream memory interfaces, prohibition of concurrent host accesses, and more). To this end, we believe that better support for fine-grained approaches is warranted in accelerators coupled with PIM-enabled memories.
A key challenge in the adoption of fine-grained PIM approaches is enforcing memory ordering. We discuss how existing memory ordering primitives (fences) are not only insufficient but their large overheads render them impractical to support fine-grain computation offloads and arbitration. To address this challenge, we make the key observation that the core-centric nature of memory ordering is unnecessary for PIM computations. We propose a novel lightweight memory ordering primitive for PIM use cases, 𝑂𝑟𝑑𝑒𝑟𝐿𝑖𝑔ℎ𝑡, which moves away from core-centric ordering enforcement and considerably reduces the overheads of enforcing correctness. For a suite of key computations from machine learning, data analytics, and genomics, we demonstrate that 𝑂𝑟𝑑𝑒𝑟𝐿𝑖𝑔ℎ𝑡 delivers 5.5× to 8.5× speedup over traditional fences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- On Consistency for Bulk-Bitwise Processing-in-MemoryBen Perach, Ronny Ronen, Shahar KvatinskyHPCA 2023 · 被引用 6 次
- CoGraf: Fully Accelerating Graph Applications with Fine-Grained PIMAli Semi Yenimol, Anirban Nag, Chang Hyun Park, David Black-SchafferASPLOS 2026
它引用的顶会 Paper1
相关 Paper
- (Almost) Fence-less Persist OrderingSara Mahdizadeh-Shahri, Seyed Armin Vakil-Ghahani, Aasheesh KolliMICRO 2020 · 被引用 15 次
- PIM-STM: Software Transactional Memory for Processing-In-Memory SystemsAndré Lopes, Daniel Castro, Paolo RomanoASPLOS 2024 · 被引用 12 次
- Accelerating Aggregation Using a Real Processing-in-Memory SystemMuhammad Attahir Jibril, Hani Al-Sayeh, Kai-Uwe SattlerICDE 2024 · 被引用 7 次
- Accelerating Transactional Execution via Processing-In-MemoryAndré Lopes, Daniel Castro, Paolo RomanoEuroSys 2026
- PIMnet: A Domain-Specific Network for Efficient Collective Communication in Scalable PIMHyojun Son, Gilbert Jonatan, Xiangyu Wu, Haeyoon Cho 等HPCA 2025 · 被引用 7 次
