PUMICE: Processing-using-Memory Integration with a Scalar Pipeline for Symbiotic Execution
Socrates S. Wong, Cecilio C. Tamarit, José F. Martínez
摘要
Existing SIMD extensions in scalar CPUs (e.g., SSE, AVX, etc.) can leverage instruction-level parallelism (ILP) because of their tight integration with the CPU pipeline. However, the vectors they employ are quite short, and this limits their ability to exploit data-level parallelism (DLP). On the other hand, processing-using-memory (PUM) accelerators are capable of exploiting massive amounts of DLP, as they typically perform computation on very long vectors (tens of thousands of elements) within the memory itself. Recent work demonstrates that orderof-magnitude speedups can be achieved by these architectures for a variety of workloads over area-equivalent multicore CPUs with SIMD extensions. Still, PUM architectures are largely decoupled from the CPU itself, thereby limiting their ability to tap the CPU's ILP the way SIMD extensions do.
In this paper, we propose PUMICE, a tightly integrated CPU-PUM architecture that simultaneously exploits DLP and ILP for very long vector operations. As a result of this tight integration, PUMICE delivers significant performance gains: Our experimental results show speedups of up to 2.2× (1.4× on average) over a state-of-the-art decoupled approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Multi-Dimensional Vector ISA Extension for Mobile In-Cache ComputingAlireza Khadem, Daichi Fujiki, Hilbert Chen, Yufeng Gu 等HPCA 2025 · 被引用 4 次
- BAAP: Coupling Compute-in-SRAM with DRAM Banks for Near-Memory ProcessingCecilio C. Tamarit, Socrates S. Wong, Akshati Vaishnav, José F. MartínezISCA 2026
它引用的顶会 Paper2
- CAPE: A Content-Addressable Processing EngineHelena Caminal, Kailin Yang, Srivatsa Srinivasa, Akshay Krishna Ramanathan 等HPCA 2021 · 被引用 30 次
- Accelerating database analytic query workloads using an associative processorHelena Caminal, Yannis Chronis, Tianshu Wu, Jignesh M. Patel 等ISCA 2022 · 被引用 19 次
相关 Paper
- The Memory Processing Unit: A Generalized Interface for End-to-End In-Memory ExecutionMinh S. Q. Truong, Yiqiu Sun, Dawei Xiong, Amol Shah 等HPCA 2026 · 被引用 1 次
- DARTH-PUM: A Hybrid Processing-Using-Memory ArchitectureRyan Wong, Ben Feinberg, Saugata GhoseASPLOS 2026 · 被引用 1 次
- Interleaved Multi-VectorizingZhuhe Fang, Beilei Zheng, Chuliang WengVLDB 2020 · 被引用 19 次
- pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup TablesJoão Dinis Ferreira, Gabriel Falcão, Juan Gómez-Luna, Mohammed Alser 等MICRO 2022 · 被引用 60 次
- RACER: Bit-Pipelined Processing Using Resistive MemoryMinh S. Q. Truong, Eric Chen, Deanyone Su, Liting Shen 等MICRO 2021 · 被引用 40 次
