Multi-Grained Multimodal Interaction Network for Entity Linking
Pengfei Luo, Tong Xu, Shiwei Wu, Chen Zhu, Linli Xu, Enhong Chen
摘要
Multimodal entity linking (MEL) task, which aims at resolving ambiguous mentions to a multimodal knowledge graph, has attracted wide attention in recent years. Though large efforts have been made to explore the complementary effect among multiple modalities, however, they may fail to fully absorb the comprehensive expression of abbreviated textual context and implicit visual indication. Even worse, the inevitable noisy data may cause inconsistency of different modalities during the learning process, which severely degenerates the performance. To address the above issues, in this paper, we propose a novel Multi-GraIned Multimodal InteraCtion Network (MIMIC) framework for solving the MEL task. Specifically, the unified inputs of mentions and entities are first encoded by textual/visual encoders separately, to extract global descriptive features and local detailed features. Then, to derive the similarity matching score for each mention-entity pair, we device three interaction units to comprehensively explore the intra-modal interaction and inter-modal fusion among features of entities and mentions. In particular, three modules, namely the Text-based Global-Local interaction Unit (TGLU), Vision-based DuaL interaction Unit (VDLU) and Cross-Modal Fusion-based interaction Unit (CMFU) are designed to capture and integrate the fine-grained representation lying in abbreviated text and implicit visual cues. Afterwards, we introduce a unit-consistency objective function via contrastive learning to avoid inconsistency and model degradation. Experimental results on three public benchmark datasets demonstrate that our solution outperforms various state-of-the-art baselines, and ablation studies verify the effectiveness of designed modules.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DRIN: Dynamic Relation Interactive Network for Multimodal Entity LinkingShangyu Xing, Fei Zhao, Zhen Wu, Chunhui Li 等ACM MM 2023 · 被引用 21 次
- RCRank: Multimodal Ranking of Root Causes of Slow Queries in Cloud Database SystemsBiao Ouyang, Yingying Zhang, Hanyin Cheng, Yang Shu 等VLDB 2025 · 被引用 7 次
- OpenMEL: Unsupervised Multimodal Entity Linking Using Noise-Free Expanded Queries and Global CoherenceXinyi Zhu, Yongqi Zhang, Lei ChenVLDB 2025 · 被引用 2 次
- I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity LinkingZiyan Liu, Junwen Li, Kaiwen Li, Tong Ruan 等ACM MM 2025 · 被引用 2 次
- Multi-level Mixture of Experts for Multimodal Entity LinkingZhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li 等KDD 2025 · 被引用 1 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- Multi-level Matching Network for Multimodal Entity LinkingZhiwei Hu, Víctor Gutiérrez-Basulto, Ru Li, Jeff Z. PanKDD 2025 · 被引用 1 次
- A Dual-Way Enhanced Framework from Text Matching Point of View for Multimodal Entity LinkingShezheng Song, Shan Zhao, Chengyu Wang, Tianwei Yan 等AAAI 2024
- Multimodal Entity Linking with Gated Hierarchical Fusion and Contrastive TrainingPeng Wang, Jiangheng Wu, Xiaohang ChenSIGIR 2022 · 被引用 52 次
- Bridging Gaps in Content and Knowledge for Multimodal Entity LinkingPengfei Luo, Tong Xu, Che Liu, Suojuan Zhang 等ACM MM 2024 · 被引用 6 次
- Multi-modal Siamese Network for Entity AlignmentLiyi Chen, Zhi Li, Tong Xu, Han Wu 等KDD 2022 · 被引用 82 次
