OpenMEL: Unsupervised Multimodal Entity Linking Using Noise-Free Expanded Queries and Global Coherence
Xinyi Zhu, Yongqi Zhang, Lei Chen
Abstract
Multimodal Entity Linking (MEL), which involves disambiguating a mention composed of multimodal inputs to a multimodal knowledge base (KB), has gained increasing attention. Although existing MEL approaches using supervised learning show promising performance, they depend heavily on large-scale labeled training data, which is expensive to obtain for each new scenario. Unsupervised learning MEL methods, on the other hand, typically consist of two main steps. In the first multimodal data encoding step, these methods either assume that the multimodal data inputs are of high quality or attempt to filter out the noisy modality. In the second entity ranking step, they employ a bipartite graph to model the relationships only between mentions and entities. However, unsupervised methods face challenges in both steps. In the first step, data quality issues arise, including limited context in textual inputs and noise in the corresponding images. Moreover, in the second step, the bipartite graph fails to capture coherence between highly correlated entities within the KB, which offers clues on shared domains among entities. This limitation hinders effective retrieval of the target entity. To address these issues, we propose a novel unsupervised learning framework, OpenMEL, for solving the MEL task. We enhance the textual modality contextual information by incorporating full context comprehension and general knowledge, and generates three levels of visual inputs for further adaptive selection to handle noise. To capture global entity coherence, we construct a tree cover structure, defining it as a maximum spanning tree with bounded nodes to meet the MEL objective. We then introduce a greedy algorithm with theoretical guarantees to solve this problem. Experimental results on three public benchmark datasets show that OpenMEL outperforms various state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 181da4ef-9c84-494c-ab59-e56fcec90433Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scalable Zero-shot Entity Linking with Dense Entity RetrievalLedell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel et al.EMNLP 2020 · 336 citations
- A Benchmarking Study of Embedding-based Entity Alignment for Knowledge GraphsZequn Sun, Qingheng Zhang, Wei Hu, Chengming Wang et al.VLDB 2020 · 297 citations
- WikiDiverse: A Multimodal Entity Linking Dataset with Diversified Contextual Topics and Entity TypesXuwu Wang, Junfeng Tian, Min Gui, Zhixu Li et al.ACL 2022 · 75 citations
Related papers
- Multimodal Entity Linking: A New Dataset and A BaselineJingru Gan, Jinchang Luo, Haiwei Wang, Shuhui Wang et al.ACM MM 2021 · 41 citations
- Multi-level Matching Network for Multimodal Entity LinkingZhiwei Hu, Víctor Gutiérrez-Basulto, Ru Li, Jeff Z. PanKDD 2025 · 1 citation
- A Dual-Way Enhanced Framework from Text Matching Point of View for Multimodal Entity LinkingShezheng Song, Shan Zhao, Chengyu Wang, Tianwei Yan et al.AAAI 2024
- Multi-level Mixture of Experts for Multimodal Entity LinkingZhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li et al.KDD 2025 · 1 citation
- Multi-Grained Multimodal Interaction Network for Entity LinkingPengfei Luo, Tong Xu, Shiwei Wu, Chen Zhu et al.KDD 2023 · 21 citations
