Multimodal Entity Linking with Gated Hierarchical Fusion and Contrastive Training
Peng Wang, Jiangheng Wu, Xiaohang Chen
Abstract
Previous entity linking methods in knowledge graphs (KGs) mostly link the textual mentions to corresponding entities. However, they have deficiencies in processing numerous multimodal data, when the text is too short to provide enough context. Consequently, we conceive the idea of introducing valuable information of other modalities, and propose a novel multimodal entity linking method with gated hierarchical multimodal fusion and contrastive training (GHMFC). Firstly, in order to discover the fine-grained inter-modal correlations, GHMFC extracts the hierarchical features of text and visual co-attention through the multi-modal co-attention mechanism: textual-guided visual attention and visual-guided textual attention. The former attention obtains weighted visual features under the guidance of textual information. In contrast, the latter attention produces weighted textual features under the guidance of visual information. Afterwards, gated fusion is used to evaluate the importance of hierarchical features of different modalities and integrate them into the final multimodal representations of mentions. Subsequently, contrastive training with two types of contrastive losses is designed to learn more generic multimodal features and reduce noise. Finally, the linking entities are selected by calculating the cosine similarity between representations of mentions and entities in KGs. To evaluate the proposed method, this paper releases two new open multimodal entity linking datasets: WikiMEL and Richpedia-MEL. Experimental results demonstrate that GHMFC can learn meaningful multimodal representation and significantly outperforms most of the baseline methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1fb30591-c92f-46f8-a762-fe5b7d0d908aCited by top-tier papers7
- DRIN: Dynamic Relation Interactive Network for Multimodal Entity LinkingShangyu Xing, Fei Zhao, Zhen Wu, Chunhui Li et al.ACM MM 2023 · 21 citations
- Multi-Grained Multimodal Interaction Network for Entity LinkingPengfei Luo, Tong Xu, Shiwei Wu, Chen Zhu et al.KDD 2023 · 21 citations
- OpenMEL: Unsupervised Multimodal Entity Linking Using Noise-Free Expanded Queries and Global CoherenceXinyi Zhu, Yongqi Zhang, Lei ChenVLDB 2025 · 2 citations
- I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity LinkingZiyan Liu, Junwen Li, Kaiwen Li, Tong Ruan et al.ACM MM 2025 · 2 citations
- Multi-level Mixture of Experts for Multimodal Entity LinkingZhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li et al.KDD 2025 · 1 citation
Related papers
- Multi-level Matching Network for Multimodal Entity LinkingZhiwei Hu, Víctor Gutiérrez-Basulto, Ru Li, Jeff Z. PanKDD 2025 · 1 citation
- A Dual-Way Enhanced Framework from Text Matching Point of View for Multimodal Entity LinkingShezheng Song, Shan Zhao, Chengyu Wang, Tianwei Yan et al.AAAI 2024
- Multi-modal Siamese Network for Entity AlignmentLiyi Chen, Zhi Li, Tong Xu, Han Wu et al.KDD 2022 · 82 citations
- HFR-MKGC: Hierarchical Fusion Reasoning with MLLMs for Multi-modal Knowledge Graph CompletionDi Wang, Junping Du, Zhe Xue, Meiyu Liang et al.AAAI 2026
- Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph CompletionYu Zhao, Ying Zhang, Baohang Zhou, Xinying Qian et al.SIGIR 2024 · 15 citations
