Data-Centric Learning from Unlabeled Graphs with Diffusion Model
Gang Liu, Eric Inae, Tong Zhao, Jiaxin Xu, Tengfei Luo, Meng Jiang
摘要
Graph property prediction tasks are important and numerous. While each task offers a small size of labeled examples, unlabeled graphs have been collected from various sources and at a large scale. A conventional approach is training a model with the unlabeled graphs on self-supervised tasks and then fine-tuning the model on the prediction tasks. However, the self-supervised task knowledge could not be aligned or sometimes conflicted with what the predictions needed. In this paper, we propose to extract the knowledge underlying the large set of unlabeled graphs as a specific set of useful data points to augment each property prediction model. We use a diffusion model to fully utilize the unlabeled graphs and design two new objectives to guide the model's denoising process with each task's labeled data to generate task-specific graph examples and their labels. Experiments demonstrate that our data-centric approach performs significantly better than fifteen existing various methods on fifteen tasks. The performance improvement brought by unlabeled data is visible as the generated labeled examples unlike the self-supervised learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Discrete-state Continuous-time Diffusion for Graph GenerationZhe Xu, Ruizhong Qiu, Yuzhong Chen, Huiyuan Chen 等NeurIPS 2024 · 被引用 92 次
- Graph Diffusion Transformers for Multi-Conditional Molecular GenerationGang Liu, Jiaxin Xu, Tengfei Luo, Meng JiangNeurIPS 2024 · 被引用 73 次
- Optimizing OOD Detection in Molecular Graphs: A Novel Approach with Diffusion ModelsXu Shen, Yili Wang, Kaixiong Zhou, Shirui Pan 等KDD 2024 · 被引用 12 次
- Cross-Domain Graph Data Scaling: A Showcase with Diffusion ModelsWenzhuo Tang, Haitao Mao, Danial Dervovic, Ivan Brugere 等NeurIPS 2025 · 被引用 8 次
- TDNetGen: Empowering Complex Network Resilience Prediction with Generative Augmentation of Topology and DynamicsChang Liu, Jingtao Ding, Yiwen Song, Yong LiKDD 2024 · 被引用 6 次
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik 等ICLR 2020 · 被引用 1,744 次
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan 等NeurIPS 2020 · 被引用 1,631 次
相关 Paper
- Graph Anomaly Detection with Few Labels: A Data-Centric ApproachXiaoxiao Ma, Ruikun Li, Fanzhen Liu, Kaize Ding 等KDD 2024 · 被引用 9 次
- KPGT: Knowledge-Guided Pre-training of Graph Transformer for Molecular Property PredictionHan Li, Dan Zhao, Jianyang ZengKDD 2022 · 被引用 55 次
- Semi-Supervised Graph Imbalanced RegressionGang Liu, Tong Zhao, Eric Inae, Tengfei Luo 等KDD 2023 · 被引用 20 次
- DisCo: Diffusion-guided Unbiased Discriminative Learning for Unsupervised Graph Domain AdaptationHaodong Zhang, Tao Ren, Changhu Wang, Yifan Wang 等KDD 2026
- Data Augmentation with Diffusion for Open-Set Semi-Supervised LearningSeonghyun Ban, Heesan Kong, Kee-Eung KimNeurIPS 2024 · 被引用 4 次
