Chunk-Wise Quantization for Graph Collaborative Filtering
Kaixi Hu, Peipei Wang, Kaize Shi, Jingling Yuan, Yu Yang, Guandong Xu, Lin Li
摘要
Energy efficiency has become a critical requirement, driving recommendation systems for resource-constrained environments such as edge devices. Model quantization offers an effective way to build low-bitwidth models while preserving accuracy. However, user–item interaction graphs contain numerous nodes and complex topological structures, leading nodes to exhibit unique similarities and differences. Existing quantization methods uniformly process parameters in high-dimensional DNN layers (e.g., linear, convolutional, or attention layers), while inadequately capturing such similarities among node embeddings. This paper proposes GraphQ, a chunk-wise quantization framework for graph collaborative filtering that supports both the training and post-training phases in a unified perspective. Our core idea is to adaptively partition node embeddings into multiple chunks based on the distribution of embedding values, and then apply chunk-wise quantization. Specifically, for quantization-aware training (QAT), we introduce learnable low-precision quantization factors that partition node embeddings into multiple chunks and are dynamically updated following message passing. For post-training quantization (PTQ), we first cluster nodes and then partition their dimensions into chunks for weight clipping. Extensive experiments on four real-world datasets show that GraphQ outperforms state-of-the-art QAT methods by an average of 27.49% in Recall@10 under the 256-dimensional embedding and 2-bit settings, and surpasses PTQ methods by 78.64% on average under 4-bit settings.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- NodeBits: A Plug-and-Play Framework for Accelerating Graph Inference by Post-Hoc Binary QuantizationQihao Cheng, Tianhao Wu, Da Yan, Haoran TangKDD 2026
- : Aggregation-Aware Quantization for Graph Neural NetworksZeyu Zhu, Fanrong Li, Zitao Mo, Qinghao Hu 等ICLR 2023
- Plug-and-Play Parameter-Efficient Tuning of Embeddings for Federated RecommendationHaochen Yuan, Yang Zhang, Xiang He, Quan Z. Sheng 等AAAI 2026
- Lightweight Embeddings for Graph Collaborative FilteringXurong Liang, Tong Chen, Lizhen Cui, Yang Wang 等SIGIR 2024 · 被引用 13 次
- Learning Binarized Graph Representations with Multi-faceted Quantization Reinforcement for Top-K RecommendationYankai Chen, Huifeng Guo, Yingxue Zhang, Chen Ma 等KDD 2022 · 被引用 27 次
