Heterformer: Transformer-based Deep Node Representation Learning on Heterogeneous Text-Rich Networks
Bowen Jin, Yu Zhang, Qi Zhu, Jiawei Han
Abstract
Representation learning on networks aims to derive a meaningful vector representation for each node, thereby facilitating downstream tasks such as link prediction, node classification, and node clustering. In heterogeneous text-rich networks, this task is more challenging due to (1) presence or absence of text: Some nodes are associated with rich textual information, while others are not; (2) diversity of types: Nodes and edges of multiple types form a heterogeneous network structure. As pretrained language models (PLMs) have demonstrated their effectiveness in obtaining widely generalizable text representations, a substantial amount of effort has been made to incorporate PLMs into representation learning on text-rich networks. However, few of them can jointly consider heterogeneous structure (network) information as well as rich textual semantic information of each node effectively. In this paper, we propose Heterformer, a Heterogeneous Network-Empowered Transformer that performs contextualized text encoding and heterogeneous structure encoding in a unified model. Specifically, we inject heterogeneous structure information into each Transformer layer when encoding node texts. Meanwhile, Heterformer is capable of characterizing node/edge type heterogeneity and encoding nodes with or without texts. We conduct comprehensive experiments on three tasks (i.e., link prediction, node classification, and node clustering) on three large-scale datasets from different domains, where Heterformer outperforms competitive baselines significantly and consistently. The code can be found at https://github.com/PeterGriffinJin/Heterformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e46c5c1-d4d6-4d3c-9aa6-a55384083e82Cited by top-tier papers7
- RAGraph: A General Retrieval-Augmented Graph Learning FrameworkXinke Jiang, Rihong Qiu, Yongxin Xu, Wentao Zhang et al.NeurIPS 2024 · 42 citations
- Patton: Language Model Pretraining on Text-Rich NetworksBowen Jin, Wentao Zhang, Yu Zhang, Yu Meng et al.ACL 2023 · 14 citations
- Unifying Text Semantics and Graph Structures for Temporal Text-attributed Graphs with Large Language ModelsSiwei Zhang, Yun Xiong, Yateng Tang, Jiarong Xu et al.NeurIPS 2025 · 9 citations
- Instruction-based Hypergraph PretrainingMingdai Yang, Zhiwei Liu, Liangwei Yang, Xiaolong Liu et al.SIGIR 2024 · 4 citations
- Can Graph Neural Networks Learn Language with Extremely Weak Text Supervision?Zihao Li, Lecheng Zheng, Bowen Jin, Dongqi Fu et al.ACL 2025
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Multi-behavior Recommendation with Graph Convolutional NetworksBowen Jin, Chen Gao, Xiangnan He, Depeng Jin et al.SIGIR 2020 · 420 citations
- Are we really making much progress?: Revisiting, benchmarking and refining heterogeneous graph neural networksQingsong Lv, Ming Ding, Qiang Liu, Yuxiang Chen et al.KDD 2021 · 249 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Edgeformers: Graph-Empowered Transformers for Representation Learning on Textual-Edge NetworksBowen Jin, Yu Zhang, Yu Meng, Jiawei HanICLR 2023 · 5 citations
- Bootstrapping Heterogeneous Graph Representation Learning via Large Language Models: A Generalized ApproachHang Gao, Chenhao Zhang, Fengge Wu, Changwen Zheng et al.AAAI 2025 · 6 citations
- GraphFormers: GNN-nested Transformers for Representation Learning on Textual GraphJunhan Yang, Zheng Liu, Shitao Xiao, Chaozhuo Li et al.NeurIPS 2021 · 262 citations
- Leveraging Contrastive Learning for Enhanced Node Representations in Tokenized Graph TransformersJinsong Chen, Hanpeng Liu, John E. Hopcroft, Kun HeNeurIPS 2024 · 23 citations
- MetaFill: Text Infilling for Meta-Path Generation on Heterogeneous Information NetworksZequn Liu, Kefei Duan, Junwei Yang, Hanwen Xu et al.EMNLP 2022 · 1 citation
