BAAP: Coupling Compute-in-SRAM with DRAM Banks for Near-Memory Processing
Cecilio C. Tamarit, Socrates S. Wong, Akshati Vaishnav, José F. Martínez
Abstract
Near-DRAM-bank logic is a form of processing-inmemory (PIM) that lowers access latency and exploits bank-level parallelism to achieve increased throughput. Major classes of memory-bound applications, including pattern matching, vector processing, and graph algorithms, stand to benefit from adding processing capabilities at the bank level. While custom banklevel PIM solutions exist for individual domains, a performant programmable solution that is able to tackle all three remains elusive, as DRAM technology imposes limitations on the complexity and scale of computing logic that can be integrated.
Recently, a number of commercial in-DRAM PIM products feature small SRAM-based scratchpads that bridge the last gap between on-die compute logic and DRAM storage. In this paper, we examine the potential and synergies of augmenting such scratchpads with Compute-in-SRAM support. Our Bank-Adjacent Associative Processor (BAAP) repurposes part of the scratchpad to provide additional in situ computing capabilities that are complementary to the existing PIM logic, primarily acting as either a SIMD unit or pattern-matching engine. We employ the commercially available UPMEM PIM architecture as a realistic foundation, although we believe that our insights are also relevant to other scratchpad-augmented PIM designs.
We evaluate BAAP using system-level, cycle-approximate, execution-driven simulations, comparing it against detailed models of commercial bank-level designs and other academic proposals, including a host-side associative processor that is also SRAM-based. Our results demonstrate that BAAP can deliver performance improvements of more than an order of magnitude across diverse workloads, including computational genomics, and the Phoenix and PrIM benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b162d233-4713-426f-a075-ec3f20cb063eBuilds on18
- Newton: A DRAM-maker's Accelerator-in-Memory (AiM) Architecture for Machine LearningMingxuan He, Choungki Song, Ilkon Kim, Chunseok Jeong et al.MICRO 2020 · 208 citations
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira et al.ASPLOS 2021 · 182 citations
- AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model InferenceJaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim et al.ASPLOS 2024 · 125 citations
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi et al.ASPLOS 2024 · 121 citations
- SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory AcceleratorXinfeng Xie, Zheng Liang, Peng Gu, Abanti Basak et al.HPCA 2021 · 111 citations
Related papers
- DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMsMingkai Chen, Tianhua Han, Cheng Liu, Shengwen Liang et al.SC 2025 · 5 citations
- AESPA: Asynchronous Execution Scheme to Exploit Bank-Level Parallelism of Processing-in-MemoryHongju Kal, Chanyoung Yoo, Won Woo RoMICRO 2023 · 19 citations
- Hyper-Ap: Enhancing Associative Processing Through A Full-Stack OptimizationYue Zha, Jing LiISCA 2020 · 31 citations
- PIM-STM: Software Transactional Memory for Processing-In-Memory SystemsAndré Lopes, Daniel Castro, Paolo RomanoASPLOS 2024 · 12 citations
- NDPBridge: Enabling Cross-Bank Coordination in Near-DRAM-Bank Processing ArchitecturesBoyu Tian, Yiwei Li, Li Jiang, Shuangyu Cai et al.ISCA 2024 · 27 citations
