BAAP: Coupling Compute-in-SRAM with DRAM Banks for Near-Memory Processing
Cecilio C. Tamarit, Socrates S. Wong, Akshati Vaishnav, José F. Martínez
摘要
Near-DRAM-bank logic is a form of processing-inmemory (PIM) that lowers access latency and exploits bank-level parallelism to achieve increased throughput. Major classes of memory-bound applications, including pattern matching, vector processing, and graph algorithms, stand to benefit from adding processing capabilities at the bank level. While custom banklevel PIM solutions exist for individual domains, a performant programmable solution that is able to tackle all three remains elusive, as DRAM technology imposes limitations on the complexity and scale of computing logic that can be integrated.
Recently, a number of commercial in-DRAM PIM products feature small SRAM-based scratchpads that bridge the last gap between on-die compute logic and DRAM storage. In this paper, we examine the potential and synergies of augmenting such scratchpads with Compute-in-SRAM support. Our Bank-Adjacent Associative Processor (BAAP) repurposes part of the scratchpad to provide additional in situ computing capabilities that are complementary to the existing PIM logic, primarily acting as either a SIMD unit or pattern-matching engine. We employ the commercially available UPMEM PIM architecture as a realistic foundation, although we believe that our insights are also relevant to other scratchpad-augmented PIM designs.
We evaluate BAAP using system-level, cycle-approximate, execution-driven simulations, comparing it against detailed models of commercial bank-level designs and other academic proposals, including a host-side associative processor that is also SRAM-based. Our results demonstrate that BAAP can deliver performance improvements of more than an order of magnitude across diverse workloads, including computational genomics, and the Phoenix and PrIM benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Newton: A DRAM-maker's Accelerator-in-Memory (AiM) Architecture for Machine LearningMingxuan He, Choungki Song, Ilkon Kim, Chunseok Jeong 等MICRO 2020 · 被引用 208 次
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira 等ASPLOS 2021 · 被引用 182 次
- AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model InferenceJaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim 等ASPLOS 2024 · 被引用 125 次
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi 等ASPLOS 2024 · 被引用 121 次
- SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory AcceleratorXinfeng Xie, Zheng Liang, Peng Gu, Abanti Basak 等HPCA 2021 · 被引用 111 次
相关 Paper
- DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMsMingkai Chen, Tianhua Han, Cheng Liu, Shengwen Liang 等SC 2025 · 被引用 5 次
- AESPA: Asynchronous Execution Scheme to Exploit Bank-Level Parallelism of Processing-in-MemoryHongju Kal, Chanyoung Yoo, Won Woo RoMICRO 2023 · 被引用 19 次
- Hyper-Ap: Enhancing Associative Processing Through A Full-Stack OptimizationYue Zha, Jing LiISCA 2020 · 被引用 31 次
- PIM-STM: Software Transactional Memory for Processing-In-Memory SystemsAndré Lopes, Daniel Castro, Paolo RomanoASPLOS 2024 · 被引用 12 次
- NDPBridge: Enabling Cross-Bank Coordination in Near-DRAM-Bank Processing ArchitecturesBoyu Tian, Yiwei Li, Li Jiang, Shuangyu Cai 等ISCA 2024 · 被引用 27 次
