Semi-Supervised Text-Attributed Graph Distillation
Yurui Lai, Samir Moustafa, Renchi Yang, Tsz Nam Chan
摘要
Text-Attributed Graphs (TAGs) have emerged as an expressive data model for integrating graph topology with rich textual semantics. Existing representation learning methods over TAGs suffer from severe scalability bottlenecks, particularly together with Large Language Models (LLMs). While data distillation offers a promising data-centric solution, existing methods fail to capture the complex interplay between graph and text modalities, struggle with the label scarcity inherent in semi-supervised settings, and lack the ability to produce the human-readable textual attributes required for downstream LLM-based tasks. To address these challenges, we propose STAD, a unified semi-supervised framework guided by the Wasserstein Distance (WSD). Grounded in our empirical findings on real TAGs, STAD introduces a graph-text collaborative encoding module that utilizes dual-pathway encoders (graph-aware and -free) within a collaborative self-training scheme to harvest reliable pseudo-labels and fuse complementary graph-text features. Furthermore, we develop a theoretically grounded WSD-based graph sketching algorithm and a cost-effective LLM text synthesis module, which leverages cluster-based keyword extraction to generate coherent, human-readable summaries for condensed nodes. Extensive experiments on benchmark datasets demonstrate that STAD achieves a state-of-the-art performance-compression trade-off in terms of both GNN- and LLM-based downstream tasks, enabling effective and efficient TAG learning or analytics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Beyond Low-frequency Information in Graph Convolutional NetworksDeyu Bo, Xiao Wang, Chuan Shi, Huawei ShenAAAI 2021 · 被引用 773 次
- Revisiting Heterophily For Graph Neural NetworksSitao Luan, Chenqing Hua, Qincheng Lu, Jiaqi Zhu 等NeurIPS 2022 · 被引用 351 次
- Graph Condensation for Graph Neural NetworksWei Jin, Lingxiao Zhao, Shichang Zhang, Yozen Liu 等ICLR 2022 · 被引用 203 次
- GraphGPT: Graph Instruction Tuning for Large Language ModelsJiabin Tang, Yuhao Yang, Wei Wei, Lei Shi 等SIGIR 2024 · 被引用 182 次
- Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation LearningXiaoxin He, Xavier Bresson, Thomas Laurent, Adam Perold 等ICLR 2024 · 被引用 151 次
相关 Paper
- Taming Language Models for Text-attributed Graph Learning with Decoupled AggregationChuang Zhou, Zhu Wang, Shengyuan Chen, Jiahe Du 等ACL 2025
- Large Language Model Meets Graph Neural Network in Knowledge DistillationShengxiang Hu, Guobing Zou, Song Yang, Shiyi Lin 等AAAI 2025 · 被引用 19 次
- SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed GraphsRuyue Liu, Rong Yin, Xiangzhen Bo, Xiaoshuai Hao 等NeurIPS 2025 · 被引用 5 次
- Quantizing Text-attributed Graphs for Semantic-Structural IntegrationJianyuan Bo, Hao Wu, Yuan FangKDD 2025
- Preference-driven Knowledge Distillation for Few-shot Node ClassificationXing Wei, Chunchun Chen, Rui Fan, Xiaofeng Cao 等NeurIPS 2025 · 被引用 2 次
