NodeBits: A Plug-and-Play Framework for Accelerating Graph Inference by Post-Hoc Binary Quantization
Qihao Cheng, Tianhao Wu, Da Yan, Haoran Tang
Abstract
Graph neural networks (GNNs) have become a leading paradigm for graph learning, powering numerous applications, such as recommender systems in e-commerce and social platforms. However, GNN inference requires expensive propagation and aggregation of node features and feature gradients, repeated across layers, the cost of which grows with the graph size. To reduce the inference cost, recent GNN-to-MLP distillation methods use vector quantization (VQ) to discretize node embeddings for MLP prediction, which introduce additional parameters such as codebook size during pre-training that need to be manually tuned to avoid accuracy drop, and which require an extra MLP training stage. In this paper, we propose NodeBits, a lightweight plug-and-play framework that accelerates graph inference via a novel post-hoc binary quantization scheme. Unlike existing GNN-to-MLP distillation methods, NodeBits decouples quantization from GNN training, so can be directly applied to any trained GNN models to accelerate their inference % with negligible accuracy drop. while maintaining competitive predictive performance. The learned node representations are discretized into compact bit strings via a data-aware binary quantizer called DaBQ, which leverages lightweight data statistics on node representations to minimize the information loss induced by binarization. During inference, NodeBits reconstructs node embeddings from these bit strings and feeds them directly to the pre-trained readout layer, avoiding training an extra MLP. To further offset information loss of binarization, we introduce NodeBits*, which fine-tunes the readout layer to better preserve model accuracy. Extensive experiments on 20 benchmark datasets demonstrate that NodeBits achieves an accuracy comparable to full-precision GNN inference while being orders of magnitude faster. NodeBits consistently outperforms existing GNN-to-MLP methods in both accuracy and inference speed.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get aaf7697f-55b3-4f25-b347-bbe76ff51360Related papers
- : Aggregation-Aware Quantization for Graph Neural NetworksZeyu Zhu, Fanrong Li, Zitao Mo, Qinghao Hu et al.ICLR 2023
- AdaGMLP: AdaBoosting GNN-to-MLP Knowledge DistillationWeigang Lu, Ziyu Guan, Wei Zhao, Yaming YangKDD 2024 · 10 citations
- VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPsLing Yang, Ye Tian, Minkai Xu, Zhongyi Liu et al.ICLR 2024 · 48 citations
- Graph-less Neural Networks: Teaching Old MLPs New Tricks Via DistillationShichang Zhang, Yozen Liu, Yizhou Sun, Neil ShahICLR 2022 · 234 citations
- Node Identifiers: Compact, Discrete Representations for Efficient Graph LearningYuankai Luo, Hongkang Li, Qijiong Liu, Lei Shi et al.ICLR 2025
