RACER: Bit-Pipelined Processing Using Resistive Memory
Minh S. Q. Truong, Eric Chen, Deanyone Su, Liting Shen, Alexander Glass, L. Richard Carley, James A. Bain, Saugata Ghose
Abstract
To combat the high energy costs of moving data between main memory and the CPU, recent works have proposed to perform processing-using-memory (PUM), a type of processing-in-memory where operations are performed on data in situ (i.e., right at the memory cells holding the data). Several common and emerging memory technologies offer the ability to perform bitwise Boolean primitive functions by having interconnected cells interact with each other, eliminating the need to use discrete CMOS compute units for several common operations. Recent PUM architectures extend upon these Boolean primitives to perform bit-serial computation using memory. Unfortunately, several practical limitations of the underlying memory devices restrict how large emerging memory arrays can be, which hinders the ability of conventional bit-serial computation approaches to deliver high performance in addition to large energy savings.
In this paper, we propose RACER, a cost-effective PUM architecture that delivers high performance and large energy savings using small arrays of resistive memories. RACER makes use of a bit-pipelining execution model, which can pipeline bit-serial w-bit computation across w small tiles. We fully design efficient control and peripheral circuitry, whose area can be amortized over small memory tiles without sacrificing memory density, and we propose an ISA abstraction for RACER to allow for easy program/compiler integration. We evaluate an implementation of RACER using NORcapable ReRAM cells across a range of microbenchmarks extracted from data-intensive applications, and find that RACER provides 107×, 12×, and 7× the performance of a 16-core CPU, a 2304-shadercore GPU, and a state-of-the-art in-SRAM compute substrate, respectively, with energy savings of 189×, 17×, and 1.3×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 665437c0-a02b-48e7-b897-25d4fb50077dCited by top-tier papers9
- Flash-Cosmos: In-Flash Bulk Bitwise Operations Using Inherent Computation Capability of NAND Flash MemoryJisung Park, Roknoddin Azizi, Geraldo F. Oliveira, Mohammad Sadrosadati et al.MICRO 2022 · 53 citations
- Venice: Improving Solid-State Drive Parallelism at Low Cost via Conflict-Free AccessesRakesh Nadig, Mohammad Sadrosadati, Haiyu Mao, Nika Mansouri-Ghiasi et al.ISCA 2023 · 29 citations
- CHOPPER: A Compiler Infrastructure for Programmable Bit-serial SIMD Processing Using Memory in DRAMXiangjun Peng, Yaohua Wang, Ming-Chang YangHPCA 2023 · 15 citations
- On Error Correction for Nonvolatile Processing-In-MemoryHüsrev Cilasun, Salonik Resch, Zamshed I. Chowdhury, Masoud Zabihi et al.ISCA 2024 · 11 citations
- On Consistency for Bulk-Bitwise Processing-in-MemoryBen Perach, Ronny Ronen, Shahar KvatinskyHPCA 2023 · 6 citations
Builds on1
Related papers
- DARTH-PUM: A Hybrid Processing-Using-Memory ArchitectureRyan Wong, Ben Feinberg, Saugata GhoseASPLOS 2026 · 1 citation
- ReCG: ReRAM-Accelerated Sparse Conjugate GradientMingjia Fan, Xiaoming Chen, Dechuang Yang, Zhou Jin et al.DAC 2024 · 5 citations
- pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup TablesJoão Dinis Ferreira, Gabriel Falcão, Juan Gómez-Luna, Mohammed Alser et al.MICRO 2022 · 60 citations
- The Memory Processing Unit: A Generalized Interface for End-to-End In-Memory ExecutionMinh S. Q. Truong, Yiqiu Sun, Dawei Xiong, Amol Shah et al.HPCA 2026 · 1 citation
- ReSiPE: ReRAM-based Single-Spiking Processing-In-Memory EngineZiru Li, Bonan Yan, Hai Helen LiDAC 2020 · 18 citations
