Text-attributed Graph Condensation via Text Selection and Attribute Matching
Haowei Han, Yuxiang Wang, Guojia Wan, Hao Wang, Shanshan Feng, Hao Huang, Jiawei Jiang, Xiao Yan
Abstract
Text-Attributed Graph (TAG) is an important type of graph structured data, where each node has a text description. TAG models usually train a Graph Neural Network (GNN) and language model jointly, which leads to high space and time consumption, especially on large datasets. To mitigate this, we propose TAGSAM, a condensation method that compresses TAGs while preserving training accuracy. TAGSAM comes with two key designs, i.e., subgraph text Selection and Attribute similarity Matching, which compress the text description and graph topology of TAG, respectively. For the texts, subgraph text selection selects and merges representative text chunks from multiple related text descriptions by maximizing mutual information. For the graph topology, popular condensation methods based on Matching Training Trajectories (MTT) suffer from high variance, which hinders accuracy. Our attribute similarity matching mitigates this issue by aligning stable similarity matrices. We evaluate TAGSAM against six state-of-the-art baselines, where it showcases superior performance. For the same compressed size, TAGSAM improves upon the best-performing baseline by an average of 4.9% in accuracy. Furthermore, it maintains competitive training accuracy even when the TAG is condensed to just 1% size. Our code is available at https://github.com/SundayVHan/TAGSAM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2c076b3-09d9-4fab-b951-4407b11db1cdBuilds on17
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Graph Condensation for Graph Neural NetworksWei Jin, Lingxiao Zhao, Shichang Zhang, Yozen Liu et al.ICLR 2022 · 203 citations
- Augmenting Low-Resource Text Classification with Graph-Grounded Pre-training and PromptingZhihao Wen, Yuan FangSIGIR 2023 · 66 citations
- GraphCLIP: Enhancing Transferability in Graph Foundation Models for Text-Attributed GraphsYun Zhu, Haizhou Shi, Xiaotang Wang, Yongchao Liu et al.WWW 2025 · 54 citations
Related papers
- Structure Balance and Gradient Matching-Based Signed Graph CondensationRong Li, Long Xu, Songbai Liu, Junkai Ji et al.AAAI 2025 · 3 citations
- Compressing LLM Knowledge into Graph Representations for Text-attributed Graphs LearningRunhuai Chen, Dian Shen, Dandan Zhang, Kaihong Huang et al.ACL 2026
- DANCE: Dynamic, Available, Neighbor-gated Condensation for Federated Text-Attributed GraphsZekai Chen, Haodong Lu, Xunkai Li, Henan Sun et al.ICML 2026
- Structure-free Graph Condensation: From Large-scale Graphs to Condensed Graph-free DataXin Zheng, Miao Zhang, Chunyang Chen, Quoc Viet Hung Nguyen et al.NeurIPS 2023 · 115 citations
- ERAlign: Energy-based Representation Alignment of GNNs and LLMs on Text-attributed GraphsXianlin Zeng, Fan Xia, Xiangyu ChenICML 2026
