Chunk-Wise Quantization for Graph Collaborative Filtering
Kaixi Hu, Peipei Wang, Kaize Shi, Jingling Yuan, Yu Yang, Guandong Xu, Lin Li
Abstract
Energy efficiency has become a critical requirement, driving recommendation systems for resource-constrained environments such as edge devices. Model quantization offers an effective way to build low-bitwidth models while preserving accuracy. However, user–item interaction graphs contain numerous nodes and complex topological structures, leading nodes to exhibit unique similarities and differences. Existing quantization methods uniformly process parameters in high-dimensional DNN layers (e.g., linear, convolutional, or attention layers), while inadequately capturing such similarities among node embeddings. This paper proposes GraphQ, a chunk-wise quantization framework for graph collaborative filtering that supports both the training and post-training phases in a unified perspective. Our core idea is to adaptively partition node embeddings into multiple chunks based on the distribution of embedding values, and then apply chunk-wise quantization. Specifically, for quantization-aware training (QAT), we introduce learnable low-precision quantization factors that partition node embeddings into multiple chunks and are dynamically updated following message passing. For post-training quantization (PTQ), we first cluster nodes and then partition their dimensions into chunks for weight clipping. Extensive experiments on four real-world datasets show that GraphQ outperforms state-of-the-art QAT methods by an average of 27.49% in Recall@10 under the 256-dimensional embedding and 2-bit settings, and surpasses PTQ methods by 78.64% on average under 4-bit settings.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- NodeBits: A Plug-and-Play Framework for Accelerating Graph Inference by Post-Hoc Binary QuantizationQihao Cheng, Tianhao Wu, Da Yan, Haoran TangKDD 2026
- : Aggregation-Aware Quantization for Graph Neural NetworksZeyu Zhu, Fanrong Li, Zitao Mo, Qinghao Hu et al.ICLR 2023
- Plug-and-Play Parameter-Efficient Tuning of Embeddings for Federated RecommendationHaochen Yuan, Yang Zhang, Xiang He, Quan Z. Sheng et al.AAAI 2026
- Lightweight Embeddings for Graph Collaborative FilteringXurong Liang, Tong Chen, Lizhen Cui, Yang Wang et al.SIGIR 2024 · 13 citations
- Learning Binarized Graph Representations with Multi-faceted Quantization Reinforcement for Top-K RecommendationYankai Chen, Huifeng Guo, Yingxue Zhang, Chen Ma et al.KDD 2022 · 27 citations
