A Dual-Way Enhanced Framework from Text Matching Point of View for Multimodal Entity Linking
Shezheng Song, Shan Zhao, Chengyu Wang, Tianwei Yan, Shasha Li, Xiaoguang Mao, Meng Wang
摘要
Multimodal Entity Linking (MEL) aims at linking ambiguous mentions with multimodal information to entity in Knowledge Graph (KG) such as Wikipedia, which plays a key role in many applications. However, existing methods suffer from shortcomings, including modality impurity such as noise in raw image and ambiguous textual entity representation, which puts obstacles to MEL. We formulate multimodal entity linking as a neural text matching problem where each multimodal information (text and image) is treated as a query, and the model learns the mapping from each query to the relevant entity from candidate entities. This paper introduces a dual-way enhanced (DWE) framework for MEL: (1) our model refines queries with multimodal data and addresses semantic gaps using cross-modal enhancers between text and image information. Besides, DWE innovatively leverages fine-grained image attributes, including facial characteristic and scene feature, to enhance and refine visual features.
(2)By using Wikipedia descriptions, DWE enriches entity semantics and obtains more comprehensive textual representation, which reduces between textual representation and the entities in KG. Extensive experiments on three public benchmarks demonstrate that our method achieves state-of-the-art (SOTA) performance, indicating the superiority of our model. The code is released on https://github.com/season1blue/DWE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- OpenViewer: Openness-Aware Multi-View LearningShide Du, Zihan Fang, Yanchao Tan, Changwei Wang 等AAAI 2025 · 被引用 5 次
- OpenMEL: Unsupervised Multimodal Entity Linking Using Noise-Free Expanded Queries and Global CoherenceXinyi Zhu, Yongqi Zhang, Lei ChenVLDB 2025 · 被引用 2 次
- I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity LinkingZiyan Liu, Junwen Li, Kaiwen Li, Tong Ruan 等ACM MM 2025 · 被引用 2 次
- Multi-level Mixture of Experts for Multimodal Entity LinkingZhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li 等KDD 2025 · 被引用 1 次
- Leveraging Image as Compressed Visual Prompt and Hierarchical Visual Knowledge for Effective Image Utilization in MLLMsShezheng Song, Kangcheng Ding, Shan Zhao, Shasha Li 等AAAI 2026
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scalable Zero-shot Entity Linking with Dense Entity RetrievalLedell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel 等EMNLP 2020 · 被引用 336 次
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng 等SIGIR 2022 · 被引用 227 次
- Mean-Shifted Contrastive Loss for Anomaly DetectionTal Reiss, Yedid HoshenAAAI 2023 · 被引用 153 次
- Neural Machine Translation with Universal Visual RepresentationZhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama 等ICLR 2020 · 被引用 117 次
相关 Paper
- Multi-level Matching Network for Multimodal Entity LinkingZhiwei Hu, Víctor Gutiérrez-Basulto, Ru Li, Jeff Z. PanKDD 2025 · 被引用 1 次
- Multimodal Entity Linking: A New Dataset and A BaselineJingru Gan, Jinchang Luo, Haiwei Wang, Shuhui Wang 等ACM MM 2021 · 被引用 41 次
- WikiDiverse: A Multimodal Entity Linking Dataset with Diversified Contextual Topics and Entity TypesXuwu Wang, Junfeng Tian, Min Gui, Zhixu Li 等ACL 2022 · 被引用 75 次
- Multi-Grained Multimodal Interaction Network for Entity LinkingPengfei Luo, Tong Xu, Shiwei Wu, Chen Zhu 等KDD 2023 · 被引用 21 次
- Multimodal Entity Linking with Gated Hierarchical Fusion and Contrastive TrainingPeng Wang, Jiangheng Wu, Xiaohang ChenSIGIR 2022 · 被引用 52 次
