FLAG: An FPGA-Based System for Low-Latency GNN Inference Service Using Vector Quantization
Yunki Han, Taehwan Kim, Jiwan Kim, Seohye Ha, Lee-Sup Kim
Abstract
Enabling real-time GNN inference services requires low end-to-end latency to meet service level agreements. However, intensive preparation steps and the neighborhood explosion problem pose significant challenges to efficient GNN inference serving. In this paper, we propose FLAG, an FPGA-based GNN inference serving system using vector quantization. To reduce preparation overhead, we introduce offline preprocessing to precompute and compress hidden embeddings for serving. A dedicated FPGA accelerator leverages the precomputed data to enable lightweight aggregation. As a result, FLAG achieves average speedups of , and on three GNN models compared to the baseline system.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get aa38c77d-38bf-4272-adf2-e706854ab0c2Cited by top-tier papers1
Ask how each one uses itRelated papers
- EOD: Enabling Low Latency GNN Inference via Near-Memory Concatenate AggregationTaehwan Kim, Yunki Han, Seohye Ha, Jiwan Kim et al.ISCA 2025 · 2 citations
- VQ-GNN: A Universal Framework to Scale up Graph Neural Networks using Vector QuantizationMucong Ding, Kezhi Kong, Jingling Li, Chen Zhu et al.NeurIPS 2021 · 68 citations
- : Aggregation-Aware Quantization for Graph Neural NetworksZeyu Zhu, Fanrong Li, Zitao Mo, Qinghao Hu et al.ICLR 2023
- FlowGNN: A Dataflow Architecture for Real-Time Workload-Agnostic Graph Neural Network InferenceRishov Sarkar, Stefan Abi-Karam, Yuqi He, Lakshmi Sathidevi et al.HPCA 2023 · 100 citations
- NodeBits: A Plug-and-Play Framework for Accelerating Graph Inference by Post-Hoc Binary QuantizationQihao Cheng, Tianhao Wu, Da Yan, Haoran TangKDD 2026
