Simple Yet Effective: Structure Guided Pre-trained Transformer for Multi-modal Knowledge Graph Reasoning
Ke Liang, Lingyuan Meng, Yue Liu, Meng Liu, Wei Wei, Suyuan Liu, Wenxuan Tu, Siwei Wang, Sihang Zhou, Xinwang Liu
Abstract
Various information in different modalities in an intuitive way in multi-modal knowledge graphs (MKGs), which are utilized in different downstream tasks, like recommendation. However, most MKGs are still far from complete, which motivates the flourishing of MKG reasoning models. Recently, with the development of general artificial intelligence, pre-trained transformers have drawn increasing attention, especially in multi-modal scenarios. However, the research of multi-modal pre-trained transformers (MPT) for knowledge graph reasoning (KGR) is still at an early stage. As the biggest difference between MKG and other multi-modal data, the rich structural information underlying the MKG is still not fully utilized in previous MPT. Most of them only use the graph structure as a retrieval map for matching images and texts connected with the same entity, which hinders their reasoning performances. To this end, the graph Structure Guided Multi-modal Pre-trained Transformer is proposed for knowledge graph reasoning (SGMPT). Specifically, the graph structure encoder is adopted for structural feature encoding. Then, a structure-guided fusion module with two simple yet effective strategies, i.e., weighted summation and alignment constraint, is designed to inject the structural information into both the textual and visual features. To the best of our knowledge, SGMPT is the first MPT for multi-modal KGR, which mines structural information underlying MKGs. Extensive experiments on FB15k-237-IMG and WN18-IMG, demonstrate that our SGMPT outperforms existing state-of-the-art models, and proves the effectiveness of the designed strategies.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1b8a0ace-115b-4486-a2ed-cc0a01baa8e4Cited by top-tier papers13
- Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual RecognitionChuanguang Yang, Xinqiang Yu, Han Yang, Zhulin An et al.AAAI 2025 · 26 citations
- Structure-Adaptive Multi-View Graph Clustering for Remote Sensing DataRenxiang Guan, Wenxuan Tu, Siwei Wang, Jiyuan Liu et al.AAAI 2025 · 26 citations
- Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity RepresentationYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu et al.AAAI 2025 · 25 citations
- Label-Free Backdoor Attacks in Vertical Federated LearningWei Shen, Wenke Huang, Guancheng Wan, Mang YeAAAI 2025 · 15 citations
- Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer NetworkXiang Fang, Wanlong Fang, Changshuo Wang, Daizong Liu et al.AAAI 2025 · 10 citations
Related papers
- Graph Reasoning Transformers for Knowledge-Aware Question AnsweringRuilin Zhao, Feng Zhao, Liang Hu, Guandong XuAAAI 2024 · 10 citations
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng et al.SIGIR 2022 · 227 citations
- MMKGR: Multi-hop Multi-modal Knowledge Graph ReasoningShangfei Zheng, Weiqing Wang, Jianfeng Qu, Hongzhi Yin et al.ICDE 2023 · 40 citations
- Multimodal Reasoning with Multimodal Knowledge GraphJunlin Lee, Yequan Wang, Jing Li, Min ZhangACL 2024 · 29 citations
- A New Pipeline for Knowledge Graph Reasoning Enhanced by Large Language Models Without Fine-TuningZhongwu Chen, Long Bai, Zixuan Li, Zhen Huang et al.EMNLP 2024 · 3 citations
