Accelerating Large Scale Real-Time GNN Inference using Channel Pruning
Hongkuan Zhou, Ajitesh Srivastava, Hanqing Zeng, Rajgopal Kannan, Viktor K. Prasanna
Abstract
Graph Neural Networks (GNNs) are proven to be powerful models to generate node embedding for downstream applications. However, due to the high computation complexity of GNN inference, it is hard to deploy GNNs for large-scale or real-time applications. In this paper, we propose to accelerate GNN inference by pruning the dimensions in each layer with negligible accuracy loss. Our pruning framework uses a novel LASSO regression formulation for GNNs to identify feature dimensions (channels) that have high influence on the output activation. We identify two inference scenarios and design pruning schemes based on their computation and memory usage for each. To further reduce the inference complexity, we effectively store and reuse hidden features of visited nodes, which significantly reduces the number of supporting nodes needed to compute the target embedding. We evaluate the proposed method with the node classification problem on five popular datasets and a real-time spam detection application. We demonstrate that the pruned GNN models greatly reduce computation and memory usage with little accuracy loss. For full inference, the proposed method achieves an average of 3.27x speedup with only 0.002 drop in F1-Micro on GPU. For batched inference, the proposed method achieves an average of 6.67x speedup with only 0.003 drop in F1-Micro on CPU. To the best of our knowledge, we are the first to accelerate large scale real-time GNN inference through channel pruning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29060250-e698-48f8-8a12-2d7f13337cb1Cited by top-tier papers22
- Graph-less Neural Networks: Teaching Old MLPs New Tricks Via DistillationShichang Zhang, Yozen Liu, Yizhou Sun, Neil ShahICLR 2022 · 234 citations
- SANCUS: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural NetworksJingshu Peng, Zhao Chen, Yingxia Shao, Yanyan Shen et al.VLDB 2022 · 76 citations
- Linkless Link Prediction via Relational DistillationZhichun Guo, William Shiao, Shichang Zhang, Yozen Liu et al.ICML 2023 · 60 citations
- VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPsLing Yang, Ye Tian, Minkai Xu, Zhongyi Liu et al.ICLR 2024 · 48 citations
- Algorithm and System Co-design for Efficient Subgraph-based Graph Representation LearningHaoteng Yin, Muhan Zhang, Yanbang Wang, Jianguo Wang et al.VLDB 2022 · 47 citations
Builds on6
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan et al.ICLR 2020 · 1,155 citations
- HyGCN: A GCN Accelerator with Hybrid ArchitectureMingyu Yan, Lei Deng, Xing Hu, Ling Liang et al.HPCA 2020 · 338 citations
- GraphNorm: A Principled Approach to Accelerating Graph Neural Network TrainingTianle Cai, Shengjie Luo, Keyulu Xu, Di He et al.ICML 2021 · 224 citations
- TinyGNN: Learning Efficient Graph Neural NetworksBencheng Yan, Chaokun Wang, Gaoyang Guo, Yunkai LouKDD 2020 · 75 citations
Related papers
- PruneGNN: Algorithm-Architecture Pruning Framework for Graph Neural Network AccelerationDeniz Gurevin, Mohsin Shan, Shaoyi Huang, Md Amit Hasan et al.HPCA 2024 · 28 citations
- FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network TrainingKezhao Huang, Haitian Jiang, Minjie Wang, Guangxuan Xiao et al.VLDB 2024 · 13 citations
- Accelerating Scalable Graph Neural Network Inference with Node-Adaptive PropagationXinyi Gao, Wentao Zhang, Junliang Yu, Yingxia Shao et al.ICDE 2024 · 15 citations
- Topology-Aware Network Pruning using Multi-stage Graph Embedding and Reinforcement LearningSixing Yu, Arya Mazaheri, Ali JannesariICML 2022 · 54 citations
- CompressGNN: Accelerating Graph Neural Network Training via Hierarchical CompressionZheng Chen, Feng Zhang, Yifei Xia, Wentao Zhang et al.KDD 2025 · 1 citation
