Kernel Ridge Regression-Based Graph Dataset Distillation
Zhe Xu, Yuzhong Chen, Menghai Pan, Huiyuan Chen, Mahashweta Das, Hao Yang, Hanghang Tong
Abstract
The huge volume of emerging graph datasets has become a double-bladed sword for graph machine learning. On the one hand, it empowers the success of a myriad of graph neural networks (GNNs) with strong empirical performance. On the other hand, training modern graph neural networks on huge graph data is computationally expensive. How to distill the given graph dataset while retaining most of the trained models' performance is a challenging problem. Existing efforts try to approach this problem by solving meta-learning-based bilevel optimization objectives. A major hurdle lies in that the exact solutions of these methods are computationally intensive and thus, most, if not all, of them are solved by approximate strategies which in turn hurt the distillation performance. In this paper, inspired by the recent advances in neural network kernel methods, we adopt a kernel ridge regression-based meta-learning objective which has a feasible exact solution. However, the computation of graph neural tangent kernel is very expensive, especially in the context of dataset distillation. As a response, we design a graph kernel, named LiteGNTK, tailored for the dataset distillation problem which is closely related to the classic random walk graph kernel. An effective model named Kernel rıdge regression-based graph Dataset Distillation (KIDD) and its variants are proposed. KIDD shows nice efficiency in both the forward and backward propagation processes. At the same time, KIDD shows strong empirical performance over 7 real-world datasets compared with the state-of-the-art distillation methods. Thanks to the ability to find the exact solution of the distillation objective, the learned training graphs by KIDD can sometimes even outperform the original whole training set with as few as 1.65% training graphs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Discrete-state Continuous-time Diffusion for Graph GenerationZhe Xu, Ruizhong Qiu, Yuzhong Chen, Huiyuan Chen et al.NeurIPS 2024 · 92 citations
- Fast Graph Condensation with Structure-based Neural Tangent KernelLin Wang, Wenqi Fan, Jiatong Li, Yao Ma et al.WWW 2024 · 45 citations
- Graph Mixup on Approximate Gromov-Wasserstein GeodesicsZhichen Zeng, Ruizhong Qiu, Zhe Xu, Zhining Liu et al.ICML 2024 · 30 citations
- Rethinking and Accelerating Graph Condensation: A Training-Free Approach with Class PartitionXinyi Gao, Guanhua Ye, Tong Chen, Wentao Zhang et al.WWW 2025 · 27 citations
- Mirage: Model-agnostic Graph Distillation for Graph ClassificationMridul Gupta, Sahil Manchanda, Hariprasad Kodamana, Sayan RanuICLR 2024 · 17 citations
Builds on14
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Graph Contrastive Learning with Adaptive AugmentationYanqiao Zhu, Yichen Xu, Feng Yu, Qiang Liu et al.WWW 2021 · 1,415 citations
- Graph Structure Learning for Robust Graph Neural NetworksWei Jin, Yao Ma, Xiaorui Liu, Xianfeng Tang et al.KDD 2020 · 604 citations
- Graph Contrastive Learning AutomatedYuning You, Tianlong Chen, Yang Shen, Zhangyang WangICML 2021 · 604 citations
Related papers
- Relational Database Distillation: From Structured Tables to Condensed Graph DataXinyi Gao, Jingxi Zhang, Lijian Chen, Tong Chen et al.WWW 2026 · 2 citations
- Efficient Graph Continual Learning via Lightweight Graph Neural Tangent Kernels-based Dataset DistillationRihong Qiu, Xinke Jiang, Yuchen Fang, Hongbin Lai et al.ICML 2025
- Self-Supervised Learning for Graph Dataset CondensationYuxiang Wang, Xiao Yan, Shiyu Jin, Hao Huang et al.KDD 2024 · 9 citations
- Simple yet Effective Graph Distillation via ClusteringYurui Lai, Taiyan Zhang, Renchi YangKDD 2025 · 1 citation
- Graph Distillation with Eigenbasis MatchingYang Liu, Deyu Bo, Chuan ShiICML 2024 · 17 citations
