Query Prior Matters: A MRC Framework for Multimodal Named Entity Recognition
Meihuizi Jia, Xin Shen, Lei Shen, Jinhui Pang, Lejian Liao, Yang Song, Meng Chen, Xiaodong He
Abstract
Multimodal named entity recognition (MNER) is a vision-language task where the system is required to detect entity spans and corresponding entity types given a sentence-image pair. Existing methods capture text-image relations with various attention mechanisms that only obtain implicit alignments between entity types and image regions. To locate regions more accurately and better model cross-/within-modal relations, we propose a machine reading comprehension based framework for MNER, namely MRC-MNER. By utilizing queries in MRC, our framework can provide prior information about entity types and image regions. Specifically, we design two stages, Query-Guided Visual Grounding and Multi-Level Modal Interaction, to align fine-grained type-region information and simulate text-image/inner-text interactions respectively. For the former, we train a visual grounding model via transfer learning to extract region candidates that can be further integrated into the second stage to enhance token representations. For the latter, we design text-image and inner-text interaction modules along with three sub-tasks for MRC-MNER. To verify the effectiveness of our model, we conduct extensive experiments on two public MNER datasets, Twitter2015 and Twitter2017. Experimental results show that MRC-MNER outperforms the current state-of-the-art models on Twitter2017, and yields competitive results on Twitter2015.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07cff759-ed65-404c-888b-3f8ee31ca4b9Cited by top-tier papers7
- MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query GroundingMeihuizi Jia, Lei Shen, Xin Shen, Lejian Liao et al.AAAI 2023 · 68 citations
- Grounded Multimodal Named Entity Recognition on Social MediaJianfei Yu, Ziyan Li, Jieming Wang, Rui XiaACL 2023 · 32 citations
- Multimodal Relation Extraction via a Mixture of Hierarchical Visual Context LearnersXiyang Liu, Chunming Hu, Richong Zhang, Kai Sun et al.WWW 2024 · 21 citations
- Learning Implicit Entity-object Relations by Bidirectional Generative Alignment for Multimodal NERFeng Chen, Jiajia Liu, Kaixiang Ji, Wang Ren et al.ACM MM 2023 · 14 citations
- Hierarchical Aligned Multimodal Learning for NER on Tweet PostsPeipei Liu, Hong Li, Yimo Ren, Jie Liu et al.AAAI 2024 · 11 citations
Builds on10
- A Unified MRC Framework for Named Entity RecognitionXiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han et al.ACL 2020 · 617 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerJianfei Yu, Jing Jiang, Li Yang, Rui XiaACL 2020 · 260 citations
- Multi-modal Graph Fusion for Named Entity Recognition with Targeted Visual GuidanceDong Zhang, Suzhong Wei, Shoushan Li, Hanqian Wu et al.AAAI 2021 · 240 citations
- Bidirectional Machine Reading Comprehension for Aspect Sentiment Triplet ExtractionShaowei Chen, Yu Wang, Jie Liu, Yuelin WangAAAI 2021 · 218 citations
Related papers
- Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity RecognitionJielong Tang, Zhenxing Wang, Ziyang Gong, Jianxing Yu et al.AAAI 2025 · 8 citations
- MCG-MNER: A Multi-Granularity Cross-Modality Generative Framework for Multimodal NER with InstructionJunjie Wu, Chen Gong, Ziqiang Cao, Guohong FuACM MM 2023 · 14 citations
- Fine-Grained Multimodal Named Entity Recognition and Grounding with a Generative FrameworkJieming Wang, Ziyan Li, Jianfei Yu, Li Yang et al.ACM MM 2023 · 11 citations
- Learning from Different text-image Pairs: A Relation-enhanced Graph Convolutional Network for Multimodal NERFei Zhao, Chunhui Li, Zhen Wu, Shangyu Xing et al.ACM MM 2022 · 59 citations
- Ground and Reconstruct: Entity-Region Bidirectional Alignment Pre-Training for Low-Resource GMNERRunwei Situ, Yi Cai, Yong Xu, Jiexin WangACM MM 2025 · 3 citations
