Is Visual Context Really Helpful for Knowledge Graph? A Representation Learning Perspective
Meng Wang, Sen Wang, Han Yang, Zheng Zhang, Xi Chen, Guilin Qi
Abstract
Visual modality recently has aroused extensive attention in the fields of knowledge graph and multimedia because a lot of real-world knowledge is multi-modal in nature. However, it is currently unclear to what extent the visual modality can improve the performance of knowledge graph tasks over unimodal models, and equally treating structural and visual features may encode too much irrelevant information from images. In this paper, we probe the utility of the auxiliary visual context from knowledge graph representation learning perspective by designing a Relation Sensitive Multi-modal Embedding model, RSME for short. RSME can automatically encourage or filter the influence of visual context during the representation learning. We also examine the effect of different visual feature encoders. Experimental results validate the superiority of our approach compared to the state-of-the-art methods. On the basis of in-depth analysis, we conclude that under appropriate circumstances models are capable of leveraging the visual input to generate better knowledge graph embeddings and vice versa.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2384ffbe-8ba1-48ce-abf8-9ed2469dceaaCited by top-tier papers23
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng et al.SIGIR 2022 · 227 citations
- Relation-enhanced Negative Sampling for Multimodal Knowledge Graph CompletionDerong Xu, Tong Xu, Shiwei Wu, Jingbo Zhou et al.ACM MM 2022 · 89 citations
- MoSE: Modality Split and Ensemble for Multimodal Knowledge Graph CompletionYu Zhao, Xiangrui Cai, Yike Wu, Haiwei Zhang et al.EMNLP 2022 · 66 citations
- LAFA: Multimodal Knowledge Graph Completion with Link Aware Fusion and AggregationBin Shang, Yinliang Zhao, Jun Liu, Di WangAAAI 2024 · 40 citations
- NativE: Multi-modal Knowledge Graph Completion in the WildYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu et al.SIGIR 2024 · 39 citations
Related papers
- VL-KGE: Vision-Language Models Meet Knowledge Graph EmbeddingsAthanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg et al.WWW 2026 · 2 citations
- REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-ExpertsXinkui Lin, Yongxiu Xu, Minghao Tang, Shilong Zhang et al.ACM MM 2025 · 2 citations
- Multi-Granularity Multi-Modal Knowledge Graph Representation Learning via Subgraph-Aware Adaptive Fusion and Hierarchical Relation ModelingPeining Li, Meiyu Liang, Wei Huang, Junping Du et al.WWW 2026
- MMKGR: Multi-hop Multi-modal Knowledge Graph ReasoningShangfei Zheng, Weiqing Wang, Jianfeng Qu, Hongzhi Yin et al.ICDE 2023 · 40 citations
- OTKGE: Multi-modal Knowledge Graph Embeddings via Optimal TransportZongsheng Cao, Qianqian Xu, Zhiyong Yang, Yuan He et al.NeurIPS 2022 · 117 citations
