Fulcrum: A Simplified Control and Access Mechanism Toward Flexible and Practical In-Situ Accelerators
Marzieh Lenjani, Patricia Gonzalez-Guerrero, Elaheh Sadredini, Shuangchen Li, Yuan Xie, Ameen Akel, Sean Eilert, Mircea R. Stan, Kevin Skadron
摘要
In-situ approaches process data very close to the memory cells, in the row buffer of each subarray. This minimizes data movement costs and affords parallelism across subarrays. However, current in-situ approaches are limited to only row-wide bitwise (or few-bit) operations applied uniformly across the row buffer. They impose a significant overhead of multiple row activations for emulating 32-bit addition and multiplications using bitwise operations and cannot support operations with data dependencies or based on predicates. Moreover, with current peripheral logic, communication among subarrays is inefficient, and with typical data layouts, bits in a word are not physically adjacent. The key insight of this work is that in-situ, single-word ALUs outperform in-situ, parallel, row-wide, bitwise ALUs by reducing the number of row activations and enabling new operations and optimizations. Our proposed lightweight access and control mechanism, Fulcrum, sequentially feeds data into the single-word ALU and enables operations with data dependencies and operations based on a predicate. For algorithms that require communication among subarrays, we augment the peripheral logic with broadcasting capabilities and a previously-proposed method for low-cost inter-subarray data movement. The sequential processor also enables overlapping of broadcasting and computation, and reuniting bits that are physically adjacent. In order to realize true subarray-level parallelism, we introduce a lightweight column-selection mechanism through shifting one-hot encoded values. This technique enables independent column selection in each subarray. We integrate Fulcrum with Compress Express Link (CXL), a new interconnect standard. Fulcrum with one memory stack delivers on average (up to) 23.4 (76) speedup over a server-class GPU, NVIDIA P100, with three stacks of HBM2 memory, (ii) 70 (228) times speedup per memory stack over the GPU, and (iii) 19 (178.9) times speedup per memory stack over an ideal model of the GPU, which only accounts for the overhead of data movement.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira 等ASPLOS 2021 · 被引用 182 次
- NDPBridge: Enabling Cross-Bank Coordination in Near-DRAM-Bank Processing ArchitecturesBoyu Tian, Yiwei Li, Li Jiang, Shuangyu Cai 等ISCA 2024 · 被引用 27 次
- Accelerating database analytic query workloads using an associative processorHelena Caminal, Yannis Chronis, Tianshu Wu, Jignesh M. Patel 等ISCA 2022 · 被引用 19 次
- CHOPPER: A Compiler Infrastructure for Programmable Bit-serial SIMD Processing Using Memory in DRAMXiangjun Peng, Yaohua Wang, Ming-Chang YangHPCA 2023 · 被引用 15 次
- Rethinking the Encoding of Integers for Scans on Skewed DataMartin Prammer, Jignesh M. PatelSIGMOD 2024 · 被引用 2 次
相关 Paper
- CXL Memory Performance for In-Memory Data ProcessingMarcel Weisgut, Daniel Ritter, Pinar Tözün, Lawrence Benson 等VLDB 2025 · 被引用 7 次
- PIPM: Partial and Incremental Page Migration for Multi-host CXL Disaggregated Shared MemoryGangqi Huang, Heiner Litz, Yuanchao XuASPLOS 2026 · 被引用 1 次
- Declarative Sub-Operators for Universal Data ProcessingMichael Jungmair, Jana GicevaVLDB 2023 · 被引用 17 次
- CorcPUM: Efficient Processing Using Cross-Point Memory via Cooperative Row-Column Access Pipelining and Adaptive Timing Optimization in SubarraysChengning Wang, Dan Feng, Wei Tong, Jingning LiuDAC 2023 · 被引用 3 次
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper 等MICRO 2023 · 被引用 133 次
