Efficient Topology-aware Data Augmentation for High-Degree Graph Neural Networks
Yurui Lai, Xiaoyang Lin, Renchi Yang, Hongtao Wang
Abstract
In recent years, graph neural networks (GNNs) have emerged as a potent tool for learning on graph-structured data and won fruitful successes in varied fields. The majority of GNNs follow the message-passing paradigm, where representations of each node are learned by recursively aggregating features of its neighbors. However, this mechanism brings severe over-smoothing and efficiency issues over high-degree graphs (HDGs), wherein most nodes have dozens (or even hundreds) of neighbors, such as social networks, transaction graphs, power grids, etc. Additionally, such graphs usually encompass rich and complex structure semantics, which are hard to capture merely by feature aggregations in GNNs.Motivated by the above limitations, we propose TADA, an efficient and effective front-mounted data augmentation framework for GNNs on HDGs. Under the hood, TADA includes two key modules: (i) feature expansion with structure embeddings, and (ii) topology- and attribute-aware graph sparsification. The former obtains augmented node features and enhanced model capacity by encoding the graph structure into high-quality structure embeddings with our highly-efficient sketching method. Further, by exploiting task-relevant features extracted from graph structures and attributes, the second module enables the accurate identification and reduction of numerous redundant/noisy edges from the input graph, thereby alleviating over-smoothing and facilitating faster feature aggregations over HDGs. Empirically, considerably improves the predictive performance of mainstream GNN models on 8 real homophilic/heterophilic HDGs in terms of node classification, while achieving efficient training and inference processes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8baba037-5d22-4959-8d6b-4c60432f8df2Cited by top-tier papers5
- A Comprehensive Benchmark on Spectral GNNs: The Impact on Efficiency, Memory, and EffectivenessNingyi Liao, Haoyu Liu, Zulun Zhu, Siqiang Luo et al.SIGMOD 2026 · 4 citations
- Simple yet Effective Graph Distillation via ClusteringYurui Lai, Taiyan Zhang, Renchi YangKDD 2025 · 1 citation
- Diffusion-Guided Graph Data AugmentationMaria Marrium, Arif Mahmood, Muhammad Haris Khan, M. Saad Shakeel et al.NeurIPS 2025 · 1 citation
- Low-Rank Few-Shot Node Classification by Node-Level Graph DiffusionYancheng Wang, Chengshuai Zhao, Dongfang Sun, huan liu et al.ICLR 2026
- Rethinking Message Passing Neural Networks with Diffusion Distance-guided Stress MajorizationHaoran Zheng, Renchi Yang, Yubo Zhou, Jianliang XuKDD 2026
Builds on38
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 1,599 citations
Related papers
- FoSR: First-order spectral rewiring for addressing oversquashing in GNNsKedar Karhadkar, Pradeep Kr. Banerjee, Guido MontúfarICLR 2023 · 7 citations
- Local Augmentation for Graph Neural NetworksSongtao Liu, Rex Ying, Hanze Dong, Lanqing Li et al.ICML 2022 · 120 citations
- NLGT: Neighborhood-based and Label-enhanced Graph Transformer Framework for Node ClassificationXiaolong Xu, Yibo Zhou, Haolong Xiang, Xiaoyong Li et al.AAAI 2025 · 5 citations
- NAFS: A Simple yet Tough-to-beat Baseline for Graph Representation LearningWentao Zhang, Zeang Sheng, Mingyu Yang, Yang Li et al.ICML 2022 · 24 citations
- When Imbalance Meets Imbalance: Structure-driven Learning for Imbalanced Graph ClassificationWei Xu, Pengkun Wang, Zhe Zhao, Binwu Wang et al.WWW 2024 · 19 citations
