Multimodal Entity Linking: A New Dataset and A Baseline
Jingru Gan, Jinchang Luo, Haiwei Wang, Shuhui Wang, Wei He, Qingming Huang
Abstract
In this paper, we introduce a new Multimodal Entity Linking (MEL) task on the multimodal data. The MEL task discovers entities in multiple modalities and various forms within large-scale multimodal data and maps multimodal mentions in a document to entities in a structured knowledge base such as Wikipedia. Different from the conventional Neural Entity Linking (NEL) task that focuses on textual information solely, MEL aims at achieving human-level disambiguation among entities in images, texts, and knowledge bases. Due to the lack of sufficient labeled data for the MEL task, we release a large-scale multimodal entity linking dataset M3EL (abbreviated for MultiModal Movie Entity Linking). Specifically, we collect reviews and images of 1,100 movies, extract textual and visual mentions, and label them with entities registered in Wikipedia. In addition, we construct a new baseline method to solve the MEL problem, which models the alignment of textual and visual mentions as a bipartite graph matching problem and solves it with an optimal-transportation-based linking method. Extensive experiments on the M3EL dataset verify the quality of the dataset and the effectiveness of the proposed method. We envision this work to be helpful for soliciting more research effort and applications regarding multimodal computing and inference in the future. We make the dataset and the baseline algorithm publicly available at https://jingrug.github.io/research/M3EL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- DRIN: Dynamic Relation Interactive Network for Multimodal Entity LinkingShangyu Xing, Fei Zhao, Zhen Wu, Chunhui Li et al.ACM MM 2023 · 21 citations
- MESED: A Multi-Modal Entity Set Expansion Dataset with Fine-Grained Semantic Classes and Hard Negative EntitiesYangning Li, Tingwei Lu, Hai-Tao Zheng, Yinghui Li et al.AAAI 2024 · 21 citations
- Multi-Grained Multimodal Interaction Network for Entity LinkingPengfei Luo, Tong Xu, Shiwei Wu, Chen Zhu et al.KDD 2023 · 21 citations
- OpenMEL: Unsupervised Multimodal Entity Linking Using Noise-Free Expanded Queries and Global CoherenceXinyi Zhu, Yongqi Zhang, Lei ChenVLDB 2025 · 2 citations
- I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity LinkingZiyan Liu, Junwen Li, Kaiwen Li, Tong Ruan et al.ACM MM 2025 · 2 citations
Builds on6
- Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerJianfei Yu, Jing Jiang, Li Yang, Rui XiaACL 2020 · 260 citations
- Graph Optimal Transport for Cross-Domain AlignmentLiqun Chen, Zhe Gan, Yu Cheng, Linjie Li et al.ICML 2020 · 193 citations
- Fine-Grained Entity Typing for Domain Independent Entity LinkingYasumasa Onoe, Greg DurrettAAAI 2020 · 94 citations
- Improving Entity Linking by Modeling Latent Entity Type InformationShuang Chen, Jinpeng Wang, Feng Jiang, Chin-Yew LinAAAI 2020 · 71 citations
- Dynamic Graph Convolutional Networks for Entity LinkingJunshuang Wu, Richong Zhang, Yongyi Mao, Hongyu Guo et al.WWW 2020 · 34 citations
Related papers
- WikiDiverse: A Multimodal Entity Linking Dataset with Diversified Contextual Topics and Entity TypesXuwu Wang, Junfeng Tian, Min Gui, Zhixu Li et al.ACL 2022 · 75 citations
- Multi-level Matching Network for Multimodal Entity LinkingZhiwei Hu, Víctor Gutiérrez-Basulto, Ru Li, Jeff Z. PanKDD 2025 · 1 citation
- A Dual-Way Enhanced Framework from Text Matching Point of View for Multimodal Entity LinkingShezheng Song, Shan Zhao, Chengyu Wang, Tianwei Yan et al.AAAI 2024
- Multi-level Mixture of Experts for Multimodal Entity LinkingZhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li et al.KDD 2025 · 1 citation
- M^3EL: A Multi-task Multi-topic Dataset for Multi-modal Entity LinkingFang Wang, Shenglin Yin, Xiaoying Bai, Minghao Hu et al.AAAI 2025 · 3 citations
