Condensing Graphs via One-Step Gradient Matching
Wei Jin, Xianfeng Tang, Haoming Jiang, Zheng Li, Danqing Zhang, Jiliang Tang, Bing Yin
Abstract
As training deep learning models on large dataset takes a lot of time and resources, it is desired to construct a small synthetic dataset with which we can train deep learning models sufficiently. There are recent works that have explored solutions on condensing image datasets through complex bi-level optimization. For instance, dataset condensation (DC) matches network gradients w.r.t. largereal data and small-synthetic data, where the network weights are optimized for multiple steps at each outer iteration. However, existing approaches have their inherent limitations: (1) they are not directly applicable to graphs where the data is discrete; and (2) the condensation process is computationally expensive due to the involved nested optimization. To bridge the gap, we investigate efficient dataset condensation tailored for graph datasets where we model the discrete graph structure as a probabilistic model. We further propose a one-step gradient matching scheme, which performs gradient matching for only one single step without training the network weights. Our theoretical analysis shows this strategy can generate synthetic graphs that lead to lower classification loss on real graphs. Extensive experiments on various graph datasets demonstrate the effectiveness and efficiency of the proposed method. In particular, we are able to reduce the dataset size by 90% while approximating up to 98% of the original performance and our method is significantly faster than multi-step gradient matching (e.g. 15× in CIFAR10 for synthesizing 500 graphs). Code is available at https://github.com/amazon-research/DosCond . CCS CONCEPTS • Computing methodologies → Neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e6f5a68-5f4d-4500-af0a-6e16b721293dCited by top-tier papers46
- Structure-free Graph Condensation: From Large-scale Graphs to Condensed Graph-free DataXin Zheng, Miao Zhang, Chunyang Chen, Quoc Viet Hung Nguyen et al.NeurIPS 2023 · 115 citations
- M3D: Dataset Condensation by Minimizing Maximum Mean DiscrepancyHansong Zhang, Shikun Li, Pengju Wang, Dan Zeng et al.AAAI 2024 · 63 citations
- Does Graph Distillation See Like Vision Dataset Counterpart?Beining Yang, Kai Wang, Qingyun Sun, Cheng Ji et al.NeurIPS 2023 · 62 citations
- Fast Graph Condensation with Structure-based Neural Tangent KernelLin Wang, Wenqi Fan, Jiatong Li, Yao Ma et al.WWW 2024 · 45 citations
- Navigating Complexity: Toward Lossless Graph Condensation via Expanding Window MatchingYuchen Zhang, Tianle Zhang, Kai Wang, Ziyao Guo et al.ICML 2024 · 38 citations
Builds on24
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan et al.ICLR 2020 · 1,155 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Graph Structure Learning for Robust Graph Neural NetworksWei Jin, Yao Ma, Xiaorui Liu, Xianfeng Tang et al.KDD 2020 · 604 citations
Related papers
- Graph Condensation for Graph Neural NetworksWei Jin, Lingxiao Zhao, Shichang Zhang, Yozen Liu et al.ICLR 2022 · 203 citations
- Improved Distribution Matching for Dataset CondensationGanlong Zhao, Guanbin Li, Yipeng Qin, Yizhou YuCVPR 2023
- Dataset Condensation via Efficient Synthetic-Data ParameterizationJang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun et al.ICML 2022 · 234 citations
- Graph Condensation for Inductive Node Representation LearningXinyi Gao, Tong Chen, Yilong Zang, Wentao Zhang et al.ICDE 2024 · 32 citations
- Bi-Directional Multi-Scale Graph Dataset Condensation via Information BottleneckXingcheng Fu, Yisen Gao, Beining Yang, Yuxuan Wu et al.AAAI 2025 · 7 citations
