FALA: Locality-Aware PIM-Host Cooperation for Graph Processing with Fine-Grained Column Access
Changmin Shin, Jaeyong Song, Seongmin Na, Jun Sung, Hongsun Jang, Jinho Lee
摘要
Graph processing is fundamental and critical to various domains, such as social networks and recommendation systems.However, its irregular memory access patterns incur significant memory bottlenecks on modern DRAM architectures, optimized for sequential access.To address this challenge, various approaches with cache and processing-in-memory (PIM) have been explored.Cache-based approaches target efficiently utilizing the cache by improving locality through graph tiling.However, a discrepancy between burst size and data makes them suffer from redundant memory accesses.PIM leverages high internal memory bandwidth but fails to exploit cache locality, showing inefficiencies in workloads with high data reuse.Hybrid approaches attempt to combine both techniques but still face limitations since they struggle with overfetching and inefficient cache utilization when vertex data has low reuse.To this end, we propose FALA, a novel hybrid architecture that synergistically integrates PIM and cache-based acceleration to overcome these limitations.FALA introduces a fine-grained column access mechanism within memory banks and a locality-aware PIM architecture to enhance memory efficiency and cache utilization.Our design minimizes unnecessary memory access while fully utilizing both high internal bandwidth and on-chip caches.We evaluate FALA against state-of-the-art cache-based, PIM, and hybrid graph processing accelerators.Experimental results demonstrate that FALA achieves 1.56× speedup and reduces energy consumption by 35.3% in geometric mean, compared to the prior art.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile DevicesYilong Zhao, Fangxin Liu, Onur Mutlu, Mingyu Gao 等ISCA 2026 · 被引用 1 次
- A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMsHongsun Jang, Jaeyong Song, Changmin Shin, Si Ung Noh 等ASPLOS 2026
- CoGraf: Fully Accelerating Graph Applications with Fine-Grained PIMAli Semi Yenimol, Anirban Nag, Chang Hyun Park, David Black-SchafferASPLOS 2026
- LoCaLUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIMJunguk Hong, Changmin Shin, Sukjin Kim, Si Ung Noh 等HPCA 2026
相关 Paper
- Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-GatherChangmin Shin, Jaeyong Song, Hongsun Jang, Dogeun Kim 等HPCA 2025 · 被引用 5 次
- Large-Scale Graph Processing on FPGAs with Caches for Thousands of Simultaneous MissesMikhail Asiatici, Paolo IenneISCA 2021 · 被引用 28 次
- ReGraph: Scaling Graph Processing on HBM-enabled FPGAs with Heterogeneous PipelinesXinyu Chen, Yao Chen, Feng Cheng, Hongshi Tan 等MICRO 2022 · 被引用 46 次
- GaaS-X: Graph Analytics Accelerator Supporting Sparse Data Representation using Crossbar ArchitecturesNagadastagiri Challapalle, Sahithi Rampalli, Linghao Song, Nandhini Chandramoorthy 等ISCA 2020 · 被引用 67 次
- GraphRing: an HMC-ring based graph processing framework with optimized data movementZerun Li, Xiaoming Chen, Yinhe HanDAC 2022 · 被引用 4 次
