CoGraf: Fully Accelerating Graph Applications with Fine-Grained PIM
Ali Semi Yenimol, Anirban Nag, Chang Hyun Park, David Black-Schaffer
Abstract
Processing-in-Memory (PIM) delivers enormous performance by taking advantage of internal DRAM bandwidth and parallelism. However, graph applications are difficult to adapt to PIM due to their irregular access patterns. We present the first Fine-Grained PIM (FGPIM) design that fully accelerates vertex-centric push-based graph applications by accelerating both their update (computing vertex updates) and apply (summing up the updates) phases.
For the update phase, we design a tuple-based LLC that can coalesce at different granularities to group graph updates together and propose multi-DRAM column processing FGPIM instructions to match the cache coalescing to the row-level parallelism of the FGPIM. With this acceleration, the apply phase becomes the bottleneck, and we propose bank-parallel FGPIM instructions with predicates to allow FGPIM to accelerate the conditional updates as well.
We achieve an average speedup in the region of interest of 1.8×/3× compared to naive FGPIM and 4.4×/9.8× compared to state-of-the-art non-PIM baseline (HBM2/DDR4), and DRAM energy reduction of 67%/86% and 88%/94%. These results show the importance of providing a complete solution that accelerates both the update and apply phases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model InferenceYufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang et al.ASPLOS 2025 · 44 citations
- AESPA: Asynchronous Execution Scheme to Exploit Bank-Level Parallelism of Processing-in-MemoryHongju Kal, Chanyoung Yoo, Won Woo RoMICRO 2023 · 19 citations
- Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache ResourcesKonstantinos Kanellopoulos, Hong Chul Nam, Nisa Bostanci, Rahul Bera et al.MICRO 2023 · 16 citations
- OrderLight: Lightweight Memory-Ordering Primitive for Efficient Fine-Grained PIM ComputationsAnirban Nag, Rajeev BalasubramonianMICRO 2021 · 7 citations
- Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-GatherChangmin Shin, Jaeyong Song, Hongsun Jang, Dogeun Kim et al.HPCA 2025 · 5 citations
Related papers
- FALA: Locality-Aware PIM-Host Cooperation for Graph Processing with Fine-Grained Column AccessChangmin Shin, Jaeyong Song, Seongmin Na, Jun Sung et al.MICRO 2025 · 5 citations
- GoPIM: GCN-Oriented Pipeline Optimization for PIM AcceleratorsSiling Yang, Shuibing He, Wenjiong Wang, Yanlong Yin et al.HPCA 2025 · 3 citations
- Accelerating Regular Path Queries over Graph Database with Processing-in-MemoryRuoyan Ma, Shengan Zheng, Guifeng Wang, Jin Pu et al.DAC 2024 · 4 citations
- GradPIM: A Practical Processing-in-DRAM Architecture for Gradient DescentHeesu Kim, Hanmin Park, Taehyun Kim, Kwanheum Cho et al.HPCA 2021 · 48 citations
- PIMGCN: A ReRAM-Based PIM Design for Graph Convolutional Network AccelerationTao Yang, Dongyue Li, Yibo Han, Yilong Zhao et al.DAC 2021 · 39 citations
