Semi-Supervised Text-Attributed Graph Distillation
Yurui Lai, Samir Moustafa, Renchi Yang, Tsz Nam Chan
Abstract
Text-Attributed Graphs (TAGs) have emerged as an expressive data model for integrating graph topology with rich textual semantics. Existing representation learning methods over TAGs suffer from severe scalability bottlenecks, particularly together with Large Language Models (LLMs). While data distillation offers a promising data-centric solution, existing methods fail to capture the complex interplay between graph and text modalities, struggle with the label scarcity inherent in semi-supervised settings, and lack the ability to produce the human-readable textual attributes required for downstream LLM-based tasks. To address these challenges, we propose STAD, a unified semi-supervised framework guided by the Wasserstein Distance (WSD). Grounded in our empirical findings on real TAGs, STAD introduces a graph-text collaborative encoding module that utilizes dual-pathway encoders (graph-aware and -free) within a collaborative self-training scheme to harvest reliable pseudo-labels and fuse complementary graph-text features. Furthermore, we develop a theoretically grounded WSD-based graph sketching algorithm and a cost-effective LLM text synthesis module, which leverages cluster-based keyword extraction to generate coherent, human-readable summaries for condensed nodes. Extensive experiments on benchmark datasets demonstrate that STAD achieves a state-of-the-art performance-compression trade-off in terms of both GNN- and LLM-based downstream tasks, enabling effective and efficient TAG learning or analytics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b8ad028-c1ae-47f6-b633-fac78be7549dBuilds on13
- Beyond Low-frequency Information in Graph Convolutional NetworksDeyu Bo, Xiao Wang, Chuan Shi, Huawei ShenAAAI 2021 · 773 citations
- Revisiting Heterophily For Graph Neural NetworksSitao Luan, Chenqing Hua, Qincheng Lu, Jiaqi Zhu et al.NeurIPS 2022 · 351 citations
- Graph Condensation for Graph Neural NetworksWei Jin, Lingxiao Zhao, Shichang Zhang, Yozen Liu et al.ICLR 2022 · 203 citations
- GraphGPT: Graph Instruction Tuning for Large Language ModelsJiabin Tang, Yuhao Yang, Wei Wei, Lei Shi et al.SIGIR 2024 · 182 citations
- Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation LearningXiaoxin He, Xavier Bresson, Thomas Laurent, Adam Perold et al.ICLR 2024 · 151 citations
Related papers
- Taming Language Models for Text-attributed Graph Learning with Decoupled AggregationChuang Zhou, Zhu Wang, Shengyuan Chen, Jiahe Du et al.ACL 2025
- Large Language Model Meets Graph Neural Network in Knowledge DistillationShengxiang Hu, Guobing Zou, Song Yang, Shiyi Lin et al.AAAI 2025 · 19 citations
- SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed GraphsRuyue Liu, Rong Yin, Xiangzhen Bo, Xiaoshuai Hao et al.NeurIPS 2025 · 5 citations
- Quantizing Text-attributed Graphs for Semantic-Structural IntegrationJianyuan Bo, Hao Wu, Yuan FangKDD 2025
- Preference-driven Knowledge Distillation for Few-shot Node ClassificationXing Wei, Chunchun Chen, Rui Fan, Xiaofeng Cao et al.NeurIPS 2025 · 2 citations
