Lune

DAC2025Top-tier venue

FLAG: An FPGA-Based System for Low-Latency GNN Inference Service Using Vector Quantization

Yunki Han, Taehwan Kim, Jiwan Kim, Seohye Ha, Lee-Sup Kim

2025Year
1Citations
1Top-tier citations

Abstract

Enabling real-time GNN inference services requires low end-to-end latency to meet service level agreements. However, intensive preparation steps and the neighborhood explosion problem pose significant challenges to efficient GNN inference serving. In this paper, we propose FLAG, an FPGA-based GNN inference serving system using vector quantization. To reduce preparation overhead, we introduce offline preprocessing to precompute and compress hidden embeddings for serving. A dedicated FPGA accelerator leverages the precomputed data to enable lightweight aggregation. As a result, FLAG achieves average speedups of 154×176×154 \times 176 \times, and 333×333 \times on three GNN models compared to the baseline system.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get aa38c77d-38bf-4272-adf2-e706854ab0c2

Cited by top-tier papers1

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines