PIM-CCA: An Efficient PIM Architecture with Optimized Integration of Configurable Functional Units
Jeehyun Kim, Donghyeon Kim, Seokwon Kang, Bongjoon Hyun, Inho Lee, Yongjun Park
Abstract
Processing-in-Memory (PIM) is a promising architecture for alleviating data movement bottlenecks by performing computations closer to memory.However, PIM workloads often encounter computational bottlenecks within the PIM itself.As these workloads become more compute-intensive by leveraging PIM's high internal bandwidth, a small set of hot code regions emerges as the primary performance bottleneck.Unfortunately, increasing the complexity of the PIM processor is difficult due to inherent memory constraints, such as area and power.Therefore, enhancing the computational capability of PIM while maintaining a lightweight design within limited silicon budgets remains highly challenging.In this paper, we propose PIM-CCA, a novel PIM architecture that integrates a Configurable Compute Accelerator (CCA) to mitigate computational bottlenecks with minimal hardware overhead.The CCA-enabled PIM design allows for the flexible configuration of compute logic, enabling acceleration across diverse workloads.The PIM-CCA compiler constructs an instruction-level dataflow graph to identify hot and compute-bound regions and offload them to the CCA.Furthermore, we analyze the interaction between the PIM threading model and resource utilization to derive the optimal thread count for efficient CCA-enabled PIM usage.We implement PIM-CCA in a cycle-accurate simulator based on a commercially available PIM system, and evaluate it using 14 representative benchmarks.The experimental results show that PIM-CCA achieves up to 1.55× performance improvement over baseline PIM systems, with only 0.036% additional area overhead, based on P&R results with limited metal layers.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f4d7e552-1f35-429c-a508-1f1374f9c8c6Cited by top-tier papers1
Ask how each one uses itRelated papers
- PIMnet: A Domain-Specific Network for Efficient Collective Communication in Scalable PIMHyojun Son, Gilbert Jonatan, Xiangyu Wu, Haeyoon Cho et al.HPCA 2025 · 7 citations
- PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM SystemsDongjae Lee, Bongjoon Hyun, Taehun Kim, Minsoo RhuMICRO 2024 · 23 citations
- Accelerating Aggregation Using a Real Processing-in-Memory SystemMuhammad Attahir Jibril, Hani Al-Sayeh, Kai-Uwe SattlerICDE 2024 · 7 citations
- Pathfinding Future PIM Architectures by Demystifying a Commercial PIM TechnologyBongjoon Hyun, Taehun Kim, Dongjae Lee, Minsoo RhuHPCA 2024 · 62 citations
- PIM-Malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) ArchitecturesDongjae Lee, Bongjoon Hyun, Youngjin Kwon, Minsoo RhuHPCA 2026 · 1 citation
