Lune

DAC2025顶会

FLAG: An FPGA-Based System for Low-Latency GNN Inference Service Using Vector Quantization

Yunki Han, Taehwan Kim, Jiwan Kim, Seohye Ha, Lee-Sup Kim

2025年份
1被引次数
1顶会引用

摘要

Enabling real-time GNN inference services requires low end-to-end latency to meet service level agreements. However, intensive preparation steps and the neighborhood explosion problem pose significant challenges to efficient GNN inference serving. In this paper, we propose FLAG, an FPGA-based GNN inference serving system using vector quantization. To reduce preparation overhead, we introduce offline preprocessing to precompute and compress hidden embeddings for serving. A dedicated FPGA accelerator leverages the precomputed data to enable lightweight aggregation. As a result, FLAG achieves average speedups of 154×176×154 \times 176 \times, and 333×333 \times on three GNN models compared to the baseline system.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get aa38c77d-38bf-4272-adf2-e706854ab0c2

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖