On Consistency for Bulk-Bitwise Processing-in-Memory
Ben Perach, Ronny Ronen, Shahar Kvatinsky
Abstract
Processing-in-memory (PIM) architectures allow software to explicitly initiate computation in the memory. This effectively makes PIM operations a new class of memory operations, alongside standard memory operations (e.g., load, store). For software correctness, it is crucial to have ordering rules for a PIM operation with other PIM operations and other memory operations, i.e., a consistency model that takes into account PIM operations is vital. To the best of our knowledge, little attention to PIM operation consistency has been given in existing works. In this paper, we focus on a specific PIM approach, named bulk-bitwise PIM. In bulk-bitwise PIM, large bitwise operations are performed directly and stored in the memory array. We show that previous solutions for the related topic of maintaining coherency of bulk-bitwise PIM have broken the host native consistency model and prevent any guaranteed correctness. As a solution, we propose and evaluate four consistency models for bulk-bitwise PIM, from strict to relaxed. Our designs also preserve coherency between PIM and the host processor. Evaluating the proposed designs’ performance with a gem5 simulation, using the YCSB short-range scan benchmark and TPC-H queries, shows that the run time overhead of guaranteeing correctness is at most 6%, and in many cases the run time is even improved. The hardware overhead of our design is less than 0.22%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e535f13b-1383-45b4-9a14-8d3ac35c1c4dBuilds on5
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks et al.ISCA 2020 · 235 citations
- TRiM: Enhancing Processor-Memory Interfaces with Scalable Tensor Reduction in MemoryJaehyun Park, Byeongho Kim, Sungmin Yun, Eojin Lee et al.MICRO 2021 · 70 citations
- RACER: Bit-Pipelined Processing Using Resistive MemoryMinh S. Q. Truong, Eric Chen, Deanyone Su, Liting Shen et al.MICRO 2021 · 40 citations
- MOUSE: Inference In Non-volatile Memory for Energy Harvesting ApplicationsSalonik Resch, S. Karen Khatamifard, Zamshed I. Chowdhury, Masoud Zabihi et al.MICRO 2020 · 37 citations
- OrderLight: Lightweight Memory-Ordering Primitive for Efficient Fine-Grained PIM ComputationsAnirban Nag, Rajeev BalasubramonianMICRO 2021 · 7 citations
Related papers
- SHERLOCK: Scheduling Efficient and Reliable Bulk Bitwise Operations in NVMsHamid Farzaneh, João Paulo C. de Lima, Ali Nezhadi Khelejani, Asif Ali Khan et al.DAC 2024 · 2 citations
- Max-PIM: Fast and Efficient Max/Min Searching in DRAMFan Zhang, Shaahin Angizi, Deliang FanDAC 2021 · 10 citations
- WISEDRAM: A Reliable Bitwise In-DRAM AcceleratorMohammad Arman Soleimani, Nezam Rohbani, Adrián Cristal Kestelman, Osman S. Unsal et al.DAC 2025 · 1 citation
- Accelerating Transactional Execution via Processing-In-MemoryAndré Lopes, Daniel Castro, Paolo RomanoEuroSys 2026
- A Case Study of Processing-in-Memory in off-the-Shelf SystemsJoel Nider, Craig Mustard, Andrada Zoltan, John Ramsden et al.USENIX ATC 2021 · 62 citations
