CoGraf: Fully Accelerating Graph Applications with Fine-Grained PIM
Ali Semi Yenimol, Anirban Nag, Chang Hyun Park, David Black-Schaffer
摘要
Processing-in-Memory (PIM) delivers enormous performance by taking advantage of internal DRAM bandwidth and parallelism. However, graph applications are difficult to adapt to PIM due to their irregular access patterns. We present the first Fine-Grained PIM (FGPIM) design that fully accelerates vertex-centric push-based graph applications by accelerating both their update (computing vertex updates) and apply (summing up the updates) phases.
For the update phase, we design a tuple-based LLC that can coalesce at different granularities to group graph updates together and propose multi-DRAM column processing FGPIM instructions to match the cache coalescing to the row-level parallelism of the FGPIM. With this acceleration, the apply phase becomes the bottleneck, and we propose bank-parallel FGPIM instructions with predicates to allow FGPIM to accelerate the conditional updates as well.
We achieve an average speedup in the region of interest of 1.8×/3× compared to naive FGPIM and 4.4×/9.8× compared to state-of-the-art non-PIM baseline (HBM2/DDR4), and DRAM energy reduction of 67%/86% and 88%/94%. These results show the importance of providing a complete solution that accelerates both the update and apply phases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model InferenceYufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang 等ASPLOS 2025 · 被引用 44 次
- AESPA: Asynchronous Execution Scheme to Exploit Bank-Level Parallelism of Processing-in-MemoryHongju Kal, Chanyoung Yoo, Won Woo RoMICRO 2023 · 被引用 19 次
- Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache ResourcesKonstantinos Kanellopoulos, Hong Chul Nam, Nisa Bostanci, Rahul Bera 等MICRO 2023 · 被引用 16 次
- OrderLight: Lightweight Memory-Ordering Primitive for Efficient Fine-Grained PIM ComputationsAnirban Nag, Rajeev BalasubramonianMICRO 2021 · 被引用 7 次
- Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-GatherChangmin Shin, Jaeyong Song, Hongsun Jang, Dogeun Kim 等HPCA 2025 · 被引用 5 次
相关 Paper
- FALA: Locality-Aware PIM-Host Cooperation for Graph Processing with Fine-Grained Column AccessChangmin Shin, Jaeyong Song, Seongmin Na, Jun Sung 等MICRO 2025 · 被引用 5 次
- GoPIM: GCN-Oriented Pipeline Optimization for PIM AcceleratorsSiling Yang, Shuibing He, Wenjiong Wang, Yanlong Yin 等HPCA 2025 · 被引用 3 次
- Accelerating Regular Path Queries over Graph Database with Processing-in-MemoryRuoyan Ma, Shengan Zheng, Guifeng Wang, Jin Pu 等DAC 2024 · 被引用 4 次
- GradPIM: A Practical Processing-in-DRAM Architecture for Gradient DescentHeesu Kim, Hanmin Park, Taehyun Kim, Kwanheum Cho 等HPCA 2021 · 被引用 48 次
- PIMGCN: A ReRAM-Based PIM Design for Graph Convolutional Network AccelerationTao Yang, Dongyue Li, Yibo Han, Yilong Zhao 等DAC 2021 · 被引用 39 次
