Unlocking Multi-Modal Potentials for Link Prediction on Dynamic Text-Attributed Graphs
Yuanyuan Xu, Wenjie Zhang, Ying Zhang, Xuemin Lin, Xiwei Xu
Abstract
Dynamic Text-Attributed Graphs (DyTAGs) are a novel graph paradigm that captures evolving temporal events (edges) alongside rich textual attributes. Existing studies can be broadly categorized into TGNN-driven and LLM-driven approaches, both of which encode textual attributes and temporal structures for DyTAG representation. We observe that DyTAGs inherently comprise three distinct modalities: temporal, textual, and structural, often exhibiting completely disjoint distributions. However, the first two modalities are largely overlooked by existing studies, leading to suboptimal performance. To address this, we propose MoMent, a multi-modal network that explicitly models, integrates, and aligns each modality to learn node representations for link prediction. Given the disjoint nature of the original modality distributions, we first construct modality-specific features and encode them using individual encoders to capture correlations across temporal patterns, semantic context, and local structures. Each encoder generates modality-specific tokens, which are then fused into comprehensive node representations with a theoretical guarantee. To avoid disjoint subspaces of these heterogeneous modalities, we propose a dual-domain alignment loss that first aligns their distributions globally and then fine-tunes coherence at the instance level. This enhances coherent representations from temporal, textual, and structural views. Extensive experiments across seven datasets show that MoMent achieves up to 17.28% accuracy improvement and up to 31x speed-up against eight baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 763dde0e-bac3-42a1-9043-645f035cf72eCited by top-tier papers1
Ask how each one uses itBuilds on31
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- EvolveGCN: Evolving Graph Convolutional Networks for Dynamic GraphsAldo Pareja, Giacomo Domeniconi, Jie Chen, Tengfei Ma et al.AAAI 2020 · 1,429 citations
- Inductive representation learning on temporal graphsDa Xu, Chuanwei Ruan, Evren Körpeoglu, Sushant Kumar et al.ICLR 2020 · 901 citations
- Inductive Representation Learning in Temporal Networks via Causal Anonymous WalksYanbang Wang, Yen-Yu Chang, Yunyu Liu, Jure Leskovec et al.ICLR 2021 · 326 citations
- Towards Better Dynamic Graph Learning: New Architecture and Unified LibraryLe Yu, Leilei Sun, Bowen Du, Weifeng LvNeurIPS 2023 · 323 citations
Related papers
- Continual-GraphLLM: Dynamic Graph Large Language Model with Invariance Regularized Adaptive Multi-Scale ExpertsTianhang Wan, Xin Wang, Haibo Chen, Longtao Huang et al.KDD 2026
- Global-Recent Semantic Reasoning on Dynamic Text-Attributed Graphs with Large Language ModelsYunan Wang, Jianxin Li, Ziwei ZhangICLR 2026 · 2 citations
- Bridging Structure and Semantics: Uncertainty-Modulated Dual-Path Diffusion for Robust Text-Attributed Graph LearningZhizhi Yu, Jiachen Liu, Qingyu Li, Dongxiao He et al.ICML 2026
- Cross-Modal Graph Attention Network for Entity AlignmentBaogui Xu, Chengjin Xu, Bing SuACM MM 2023 · 21 citations
- ERAlign: Energy-based Representation Alignment of GNNs and LLMs on Text-attributed GraphsXianlin Zeng, Fan Xia, Xiangyu ChenICML 2026
