Self-Supervised Learning for Graph Dataset Condensation
Yuxiang Wang, Xiao Yan, Shiyu Jin, Hao Huang, Quanqing Xu, Qingchen Zhang, Bo Du, Jiawei Jiang
摘要
Graph dataset condensation (GDC) reduces a dataset with many graphs into a smaller dataset with fewer graphs while maintaining model training accuracy. GDC saves the storage cost and hence accelerates training. Although several GDC methods have been proposed, they are all supervised and require massive labels for the graphs, while graph labels can be scarce in many practical scenarios. To fill this gap, we propose a self-supervised graph dataset condensation method called SGDC, which does not require label information. Our initial design starts with the classical bilevel optimization paradigm for dataset condensation and incorporates contrastive learning techniques. But such a solution yields poor accuracy due to the biased gradient estimation caused by data augmentation. To solve this problem, we introduce representation matching, which conducts training by aligning the representations produced by the condensed graphs with the target representations generated by a pre-trained SSL model. This design eliminates the need for data augmentation and avoids biased gradient. We further propose a graph attention kernel, which not only improves accuracy but also reduces running time when combined with self-supervised kernel ridge regression (KRR). To simplify SGDC and make it more robust, we adopt a adjacency matrix reusing approach, which reuses the topology of the original graphs for the condensed graphs instead of repeatedly learning topology during training. Our evaluations on seven graph datasets find that SGDC improves model accuracy by up to 9.7% compared with 5 state-of-the-art baselines, even if they use label information. Moreover, SGDC is significantly more efficient than the baselines.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Rethinking and Accelerating Graph Condensation: A Training-Free Approach with Class PartitionXinyi Gao, Guanhua Ye, Tong Chen, Wentao Zhang 等WWW 2025 · 被引用 27 次
- Towards Pre-trained Graph Condensation via Optimal TransportYeyu Yan, Shuai Zheng, Wenjun Hui, Xiangkai Zhu 等NeurIPS 2025 · 被引用 3 次
- Contrastive Graph Condensation: Advancing Data Versatility through Self-Supervised LearningXinyi Gao, Yayong Li, Tong Chen, Guanhua Ye 等KDD 2025 · 被引用 1 次
- Text-attributed Graph Condensation via Text Selection and Attribute MatchingHaowei Han, Yuxiang Wang, Guojia Wan, Hao Wang 等WWW 2026
相关 Paper
- ST-GCond: Self-supervised and Transferable Graph Dataset CondensationBeining Yang, Qingyun Sun, Cheng Ji, Xingcheng Fu 等ICLR 2025
- Fast Graph Condensation with Structure-based Neural Tangent KernelLin Wang, Wenqi Fan, Jiatong Li, Yao Ma 等WWW 2024 · 被引用 45 次
- Kernel Ridge Regression-Based Graph Dataset DistillationZhe Xu, Yuzhong Chen, Menghai Pan, Huiyuan Chen 等KDD 2023 · 被引用 37 次
- Rethinking and Scaling Up Graph Contrastive Learning: An Extremely Efficient Approach with Group DiscriminationYizhen Zheng, Shirui Pan, Vincent C. S. Lee, Yu Zheng 等NeurIPS 2022 · 被引用 153 次
- Structure Balance and Gradient Matching-Based Signed Graph CondensationRong Li, Long Xu, Songbai Liu, Junkai Ji 等AAAI 2025 · 被引用 3 次
