Lune

MICRO2025顶会

FALA: Locality-Aware PIM-Host Cooperation for Graph Processing with Fine-Grained Column Access

Changmin Shin, Jaeyong Song, Seongmin Na, Jun Sung, Hongsun Jang, Jinho Lee

2025年份
5被引次数
4顶会引用

摘要

Graph processing is fundamental and critical to various domains, such as social networks and recommendation systems.However, its irregular memory access patterns incur significant memory bottlenecks on modern DRAM architectures, optimized for sequential access.To address this challenge, various approaches with cache and processing-in-memory (PIM) have been explored.Cache-based approaches target efficiently utilizing the cache by improving locality through graph tiling.However, a discrepancy between burst size and data makes them suffer from redundant memory accesses.PIM leverages high internal memory bandwidth but fails to exploit cache locality, showing inefficiencies in workloads with high data reuse.Hybrid approaches attempt to combine both techniques but still face limitations since they struggle with overfetching and inefficient cache utilization when vertex data has low reuse.To this end, we propose FALA, a novel hybrid architecture that synergistically integrates PIM and cache-based acceleration to overcome these limitations.FALA introduces a fine-grained column access mechanism within memory banks and a locality-aware PIM architecture to enhance memory efficiency and cache utilization.Our design minimizes unnecessary memory access while fully utilizing both high internal bandwidth and on-chip caches.We evaluate FALA against state-of-the-art cache-based, PIM, and hybrid graph processing accelerators.Experimental results demonstrate that FALA achieves 1.56× speedup and reduces energy consumption by 35.3% in geometric mean, compared to the prior art.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get a3d70fc3-e203-48a4-bbef-81bf98913164

引用它的顶会 Paper4

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖