Harnessing Language Model for Cross-Heterogeneity Graph Knowledge Transfer
Jinyu Yang, Ruijia Wang, Cheng Yang, Bo Yan, Qimin Zhou, Yang Juan, Chuan Shi
Abstract
Heterogeneous graphs (HGs) that contain various node and edge types are ubiquitous in real-world scenarios. Considering the common label sparsity problem in HGs, some researchers propose to pretrain on source HGs to extract general knowledge and then fine-tune on a target HG for knowledge transfer. However, existing methods often assume that source and target HGs share a single heterogeneity, meaning that they have the same types of nodes and edges, which contradicts the real-world scenarios requiring cross-heterogeneity transfer. Although a recent study has made some preliminary attempts in cross-heterogeneity learning, its definition of general knowledge heavily rely on human knowledge, which lacks flexibility and further leads to a suboptimal transfer. To address the problem, we propose a novel Language Model-enhanced Cross-Heterogeneity learning model, namely LMCH. Specifically, we first design a metapath-based corpus construction method to unify HG representations as languages. The corpora of source HGs are then used to fine-tune a pretrained Language Model (LM), enabling the LM to autonomously extract general knowledge across different HGs. Furthermore, to fully utilize the extensive unlabeled nodes in a few-labeled target HG, we propose an iterative training pipeline with the help of an extra Graph Neural Network (GNN) predictor, enhanced by LM-GNN contrastive alignment at the end of each iteration. Extensive experiments on four real-world datasets have demonstrated the superior performance of LMCH over state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 009d8315-abac-494d-adcb-b1ddc28c1f19Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- MAGNN: Metapath Aggregated Graph Neural Network for Heterogeneous Graph EmbeddingXinyu Fu, Jiani Zhang, Ziqiao Meng, Irwin KingWWW 2020 · 1,149 citations
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 463 citations
- Graph Meta Learning via Local SubgraphsKexin Huang, Marinka ZitnikNeurIPS 2020 · 205 citations
Related papers
- MUG: Meta-path-aware Universal Heterogeneous Graph Pre-TrainingLianze Shan, Jitao Zhao, Dongxiao He, Yongqi Huang et al.AAAI 2026 · 1 citation
- Bootstrapping Heterogeneous Graph Representation Learning via Large Language Models: A Generalized ApproachHang Gao, Chenhao Zhang, Fengge Wu, Changwen Zheng et al.AAAI 2025 · 6 citations
- Pre-training on Large-Scale Heterogeneous GraphXunqiang Jiang, Tianrui Jia, Yuan Fang, Chuan Shi et al.KDD 2021 · 44 citations
- Scalable Multi-Source Pre-training for Graph Neural NetworksMingkai Lin, Wenzhong Li, Xiaobin Hong, Sanglu LuACM MM 2024 · 2 citations
- HGACLLM: Attribute Completion in Heterogeneous Graph with Integration of External Knowledge from Large Language ModelsZongxing Zhao, Shenzhi Yang, Xingkai Yao, Yuying Wang et al.ACM MM 2025 · 1 citation
