OrderLight: Lightweight Memory-Ordering Primitive for Efficient Fine-Grained PIM Computations
Anirban Nag, Rajeev Balasubramonian
Abstract
Modern workloads such as neural networks, genomic analysis, and data analytics exhibit significant data-intensive phases (low compute to byte ratio) and, as such, stand to gain considerably by using processing-in-memory (PIM) solutions along with more traditional accelerators. While PIM has been researched extensively, the granularity of computation offload to PIM and the granularity of memory access arbitration between host and PIM, as well as their implications, have received relatively little attention. In this work, we first introduce a taxonomy to study the design space whilst considering these two aspects. Based on this taxonomy, we observe that much of PIM research to date has largely relied on coarse-grained approaches which, we argue, have steep costs (incompatibility with mainstream memory interfaces, prohibition of concurrent host accesses, and more). To this end, we believe that better support for fine-grained approaches is warranted in accelerators coupled with PIM-enabled memories.
A key challenge in the adoption of fine-grained PIM approaches is enforcing memory ordering. We discuss how existing memory ordering primitives (fences) are not only insufficient but their large overheads render them impractical to support fine-grain computation offloads and arbitration. To address this challenge, we make the key observation that the core-centric nature of memory ordering is unnecessary for PIM computations. We propose a novel lightweight memory ordering primitive for PIM use cases, ๐๐๐๐๐๐ฟ๐๐โ๐ก, which moves away from core-centric ordering enforcement and considerably reduces the overheads of enforcing correctness. For a suite of key computations from machine learning, data analytics, and genomics, we demonstrate that ๐๐๐๐๐๐ฟ๐๐โ๐ก delivers 5.5ร to 8.5ร speedup over traditional fences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 659e2298-9f2e-4434-a59d-ffae716819deCited by top-tier papers2
- On Consistency for Bulk-Bitwise Processing-in-MemoryBen Perach, Ronny Ronen, Shahar KvatinskyHPCA 2023 ยท 6 citations
- CoGraf: Fully Accelerating Graph Applications with Fine-Grained PIMAli Semi Yenimol, Anirban Nag, Chang Hyun Park, David Black-SchafferASPLOS 2026
Builds on1
Related papers
- (Almost) Fence-less Persist OrderingSara Mahdizadeh-Shahri, Seyed Armin Vakil-Ghahani, Aasheesh KolliMICRO 2020 ยท 15 citations
- PIM-STM: Software Transactional Memory for Processing-In-Memory SystemsAndrรฉ Lopes, Daniel Castro, Paolo RomanoASPLOS 2024 ยท 12 citations
- Accelerating Aggregation Using a Real Processing-in-Memory SystemMuhammad Attahir Jibril, Hani Al-Sayeh, Kai-Uwe SattlerICDE 2024 ยท 7 citations
- Accelerating Transactional Execution via Processing-In-MemoryAndrรฉ Lopes, Daniel Castro, Paolo RomanoEuroSys 2026
- PIMnet: A Domain-Specific Network for Efficient Collective Communication in Scalable PIMHyojun Son, Gilbert Jonatan, Xiangyu Wu, Haeyoon Cho et al.HPCA 2025 ยท 7 citations
