SaVe-TAG: LLM-based Interpolation for Long-Tailed Text-Attributed Graphs
Leyao Wang, Yu Wang, Bo Ni, Yuying Zhao, Hanyu Wang, Yao Ma, Tyler Derr
Abstract
Real-world graph data often follows long-tailed distributions, making it difficult for Graph Neural Networks (GNNs) to generalize well across both head and tail classes. Recent advances in Vicinal Risk Minimization (VRM) have shown promise in mitigating class imbalance with numeric interpolation; however, existing approaches largely rely on embedding-space arithmetic, which fails to capture the rich semantics inherent in text-attributed graphs. In this work, we propose our method SaVe-TAG (Semantic-aware Vicinal Risk Minimization for Long-Tailed Text-Attributed Graphs), a novel VRM framework that leverages Large Language Models (LLMs) to perform text-level interpolation, generating on-manifold, boundary-enriching synthetic samples for minority classes. To mitigate the risk of noisy generation, we introduce a confidence-based edge assignment mechanism that uses graph topology as a natural filter to ensure structural consistency. We provide theoretical justification for our method and conduct extensive experiments on benchmark datasets, showing that our approach consistently outperforms both numeric interpolation and prior long-tailed node classification baselines. Our results highlight the importance of integrating semantic and structural signals for balanced and effective learning on text-attributed graphs. The source code is publicly available at: https://github.com/LWang-Laura/SaVe-TAG . CCS Concepts • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba168665-b59a-4df0-8701-16059ba2694dCited by top-tier papers1
Ask how each one uses itBuilds on7
- Mixup for Node and Graph ClassificationYiwei Wang, Wei Wang, Yuxuan Liang, Yujun Cai et al.WWW 2021 · 220 citations
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming YinEMNLP 2023 · 102 citations
- GraphSR: A Data Augmentation Algorithm for Imbalanced Node ClassificationMengting Zhou, Zhiguo GongAAAI 2023 · 48 citations
- Hyperbolic Geometric Graph Representation Learning for Hierarchy-imbalance Node ClassificationXingcheng Fu, Yuecen Wei, Qingyun Sun, Haonan Yuan et al.WWW 2023 · 41 citations
- Leveraging Large Language Models for Node Generation in Few-Shot Learning on Text-Attributed GraphsJianxiang Yu, Yuxiang Ren, Chenghua Gong, Jiaqi Tan et al.AAAI 2025 · 32 citations
Related papers
- Test-Time Training on Graphs with Large Language Models (LLMs)Jiaxin Zhang, Yiqi Wang, Xihong Yang, Siwei Wang et al.ACM MM 2024 · 6 citations
- Generalization Principles for Inference over Text-Attributed Graphs with Large Language ModelsHaoyu Peter Wang, Shikun Liu, Rongzhe Wei, Pan LiICML 2025
- Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation LearningXiaoxin He, Xavier Bresson, Thomas Laurent, Adam Perold et al.ICLR 2024 · 151 citations
- UTAG: Leveraging LLM as a Unified Embedding Generator for Text-Attributed GraphsMingqian Ding, Jianjun Li, Zhiyuan Ma, Liwei Zhang et al.WWW 2026
- GAugLLM: Improving Graph Contrastive Learning for Text-Attributed Graphs with Large Language ModelsYi Fang, Dongzhe Fan, Daochen Zha, Qiaoyu TanKDD 2024 · 20 citations
