Cross-Domain Graph Data Scaling: A Showcase with Diffusion Models
Wenzhuo Tang, Haitao Mao, Danial Dervovic, Ivan Brugere, Saumitra Mishra, Yuying Xie, Jiliang Tang
摘要
Models for natural language and images benefit from data scaling behavior: the more data fed into the model, the better they perform. This 'better with more' phenomenon enables the effectiveness of large-scale pre-training on vast amounts of data. However, current graph pre-training methods struggle to scale up data due to heterogeneity across graphs. To achieve effective data scaling, we aim to develop a general model that is able to capture diverse data patterns of graphs and can be utilized to adaptively help the downstream tasks. To this end, we propose UniAug, a universal graph structure augmentor built on a diffusion model. We first pre-train a discrete diffusion model on thousands of graphs across domains to learn the graph structural patterns. In the downstream phase, we provide adaptive enhancement by conducting graph structure augmentation with the help of the pre-trained diffusion model via guided generation. By leveraging the pre-trained diffusion model for structure augmentation, we consistently achieve performance improvements across various downstream tasks in a plug-and-play manner. To the best of our knowledge, this study represents the first demonstration of a data-scaling graph structure augmentor on graphs across domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- GRAVER: Generative Graph Vocabularies for Robust Graph Foundation Models Fine-tuningHaonan Yuan, Qingyun Sun, Junhua Shi, Xingcheng Fu 等NeurIPS 2025 · 被引用 17 次
- How Much Can Transfer? BRIDGE: Bounded Multi-Domain Graph Foundation Model with Generalization GuaranteesHaonan Yuan, Qingyun Sun, Junhua Shi, Xingcheng Fu 等ICML 2025
它引用的顶会 Paper60
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Data-Centric Learning from Unlabeled Graphs with Diffusion ModelGang Liu, Eric Inae, Tong Zhao, Jiaxin Xu 等NeurIPS 2023 · 被引用 32 次
- MUG: Meta-path-aware Universal Heterogeneous Graph Pre-TrainingLianze Shan, Jitao Zhao, Dongxiao He, Yongqi Huang 等AAAI 2026 · 被引用 1 次
- LEDA: Latent Semantic Distribution Alignment for Multi-domain Graph Pre-trainingLianze Shan, Jitao Zhao, Dongxiao He, Siqi Liu 等WWW 2026
- Handling Feature Heterogeneity with Learnable Graph PatchesYifei Sun, Yang Yang, Xiao Feng, Zijun Wang 等KDD 2025 · 被引用 1 次
- UTAG: Leveraging LLM as a Unified Embedding Generator for Text-Attributed GraphsMingqian Ding, Jianjun Li, Zhiyuan Ma, Liwei Zhang 等WWW 2026
