FALA: Locality-Aware PIM-Host Cooperation for Graph Processing with Fine-Grained Column Access
Changmin Shin, Jaeyong Song, Seongmin Na, Jun Sung, Hongsun Jang, Jinho Lee
Abstract
Graph processing is fundamental and critical to various domains, such as social networks and recommendation systems.However, its irregular memory access patterns incur significant memory bottlenecks on modern DRAM architectures, optimized for sequential access.To address this challenge, various approaches with cache and processing-in-memory (PIM) have been explored.Cache-based approaches target efficiently utilizing the cache by improving locality through graph tiling.However, a discrepancy between burst size and data makes them suffer from redundant memory accesses.PIM leverages high internal memory bandwidth but fails to exploit cache locality, showing inefficiencies in workloads with high data reuse.Hybrid approaches attempt to combine both techniques but still face limitations since they struggle with overfetching and inefficient cache utilization when vertex data has low reuse.To this end, we propose FALA, a novel hybrid architecture that synergistically integrates PIM and cache-based acceleration to overcome these limitations.FALA introduces a fine-grained column access mechanism within memory banks and a locality-aware PIM architecture to enhance memory efficiency and cache utilization.Our design minimizes unnecessary memory access while fully utilizing both high internal bandwidth and on-chip caches.We evaluate FALA against state-of-the-art cache-based, PIM, and hybrid graph processing accelerators.Experimental results demonstrate that FALA achieves 1.56× speedup and reduces energy consumption by 35.3% in geometric mean, compared to the prior art.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a3d70fc3-e203-48a4-bbef-81bf98913164Cited by top-tier papers4
- COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile DevicesYilong Zhao, Fangxin Liu, Onur Mutlu, Mingyu Gao et al.ISCA 2026 · 1 citation
- A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMsHongsun Jang, Jaeyong Song, Changmin Shin, Si Ung Noh et al.ASPLOS 2026
- CoGraf: Fully Accelerating Graph Applications with Fine-Grained PIMAli Semi Yenimol, Anirban Nag, Chang Hyun Park, David Black-SchafferASPLOS 2026
- LoCaLUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIMJunguk Hong, Changmin Shin, Sukjin Kim, Si Ung Noh et al.HPCA 2026
Related papers
- Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-GatherChangmin Shin, Jaeyong Song, Hongsun Jang, Dogeun Kim et al.HPCA 2025 · 5 citations
- Large-Scale Graph Processing on FPGAs with Caches for Thousands of Simultaneous MissesMikhail Asiatici, Paolo IenneISCA 2021 · 28 citations
- ReGraph: Scaling Graph Processing on HBM-enabled FPGAs with Heterogeneous PipelinesXinyu Chen, Yao Chen, Feng Cheng, Hongshi Tan et al.MICRO 2022 · 46 citations
- GaaS-X: Graph Analytics Accelerator Supporting Sparse Data Representation using Crossbar ArchitecturesNagadastagiri Challapalle, Sahithi Rampalli, Linghao Song, Nandhini Chandramoorthy et al.ISCA 2020 · 67 citations
- GraphRing: an HMC-ring based graph processing framework with optimized data movementZerun Li, Xiaoming Chen, Yinhe HanDAC 2022 · 4 citations
