Self-Supervised Learning for Graph Dataset Condensation
Yuxiang Wang, Xiao Yan, Shiyu Jin, Hao Huang, Quanqing Xu, Qingchen Zhang, Bo Du, Jiawei Jiang
Abstract
Graph dataset condensation (GDC) reduces a dataset with many graphs into a smaller dataset with fewer graphs while maintaining model training accuracy. GDC saves the storage cost and hence accelerates training. Although several GDC methods have been proposed, they are all supervised and require massive labels for the graphs, while graph labels can be scarce in many practical scenarios. To fill this gap, we propose a self-supervised graph dataset condensation method called SGDC, which does not require label information. Our initial design starts with the classical bilevel optimization paradigm for dataset condensation and incorporates contrastive learning techniques. But such a solution yields poor accuracy due to the biased gradient estimation caused by data augmentation. To solve this problem, we introduce representation matching, which conducts training by aligning the representations produced by the condensed graphs with the target representations generated by a pre-trained SSL model. This design eliminates the need for data augmentation and avoids biased gradient. We further propose a graph attention kernel, which not only improves accuracy but also reduces running time when combined with self-supervised kernel ridge regression (KRR). To simplify SGDC and make it more robust, we adopt a adjacency matrix reusing approach, which reuses the topology of the original graphs for the condensed graphs instead of repeatedly learning topology during training. Our evaluations on seven graph datasets find that SGDC improves model accuracy by up to 9.7% compared with 5 state-of-the-art baselines, even if they use label information. Moreover, SGDC is significantly more efficient than the baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- Rethinking and Accelerating Graph Condensation: A Training-Free Approach with Class PartitionXinyi Gao, Guanhua Ye, Tong Chen, Wentao Zhang et al.WWW 2025 · 27 citations
- Towards Pre-trained Graph Condensation via Optimal TransportYeyu Yan, Shuai Zheng, Wenjun Hui, Xiangkai Zhu et al.NeurIPS 2025 · 3 citations
- Contrastive Graph Condensation: Advancing Data Versatility through Self-Supervised LearningXinyi Gao, Yayong Li, Tong Chen, Guanhua Ye et al.KDD 2025 · 1 citation
- Text-attributed Graph Condensation via Text Selection and Attribute MatchingHaowei Han, Yuxiang Wang, Guojia Wan, Hao Wang et al.WWW 2026
Related papers
- ST-GCond: Self-supervised and Transferable Graph Dataset CondensationBeining Yang, Qingyun Sun, Cheng Ji, Xingcheng Fu et al.ICLR 2025
- Fast Graph Condensation with Structure-based Neural Tangent KernelLin Wang, Wenqi Fan, Jiatong Li, Yao Ma et al.WWW 2024 · 45 citations
- Kernel Ridge Regression-Based Graph Dataset DistillationZhe Xu, Yuzhong Chen, Menghai Pan, Huiyuan Chen et al.KDD 2023 · 37 citations
- Rethinking and Scaling Up Graph Contrastive Learning: An Extremely Efficient Approach with Group DiscriminationYizhen Zheng, Shirui Pan, Vincent C. S. Lee, Yu Zheng et al.NeurIPS 2022 · 153 citations
- Structure Balance and Gradient Matching-Based Signed Graph CondensationRong Li, Long Xu, Songbai Liu, Junkai Ji et al.AAAI 2025 · 3 citations
