Fine-Grained Multimodal Named Entity Recognition and Grounding with a Generative Framework
Jieming Wang, Ziyan Li, Jianfei Yu, Li Yang, Rui Xia
Abstract
Multimodal Named Entity Recognition (MNER) aims to locate and classify named entities mentioned in a pair of text and image. However, most previous MNER works focus on extracting entities in the form of text but failing to ground text symbols to their corresponding visual objects. Moreover, existing MNER studies primarily classify entities into four coarse-grained entity types, which are often insufficient to map them to their real-world referents. To solve these limitations, we introduce a task named Fine-grained Multimodal Named Entity Recognition and Grounding (FMNERG) in this paper, which aims to simultaneously extract named entities in text, their fine-grained entity types, and their grounded visual objects in image. Moreover, we construct a Twitter dataset for the FMNERG task, and further propose a T5-based multImodal GEneration fRamework (TIGER), which formulates FMNERG as a generation problem by converting all the entity-type-object triples into a target sequence and adapts a pre-trained sequence-to-sequence model T5 to directly generate the target sequence from an image-text input pair. Experimental results demonstrate that TIGER performs significantly better than a number of baseline systems on the annotated Twitter dataset. Our dataset annotation and source code are publicly released at https://github.com/NUSTM/FMNERG.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b032e689-201c-4032-978a-c567072681e0Cited by top-tier papers5
- Cross-modal Multi-task Learning for Multimedia Event ExtractionJianwei Cao, Yanli Hu, Zhen Tan, Xiang ZhaoAAAI 2025 · 8 citations
- Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity RecognitionJielong Tang, Zhenxing Wang, Ziyang Gong, Jianxing Yu et al.AAAI 2025 · 8 citations
- SAKE: Self-aware Knowledge Exploitation-Exploration for Grounded Multimodal Named Entity RecognitionJielong Tang, Xujie Yuan, Jiayang Liu, Jianxing Yu et al.KDD 2026 · 1 citation
- ISR: Self-Refining Referring Expressions for Entity GroundingZhuocheng Yu, Bingchan Zhao, Yifan Song, Sujian Li et al.ACL 2025 · 1 citation
- MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity RecognitionXinkui Lin, Yuhui Zhang, Yongxiu Xu, Kun Huang et al.EMNLP 2025
Related papers
- Grounded Multimodal Named Entity Recognition on Social MediaJianfei Yu, Ziyan Li, Jieming Wang, Rui XiaACL 2023 · 32 citations
- MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query GroundingMeihuizi Jia, Lei Shen, Xin Shen, Lejian Liao et al.AAAI 2023 · 68 citations
- Query Prior Matters: A MRC Framework for Multimodal Named Entity RecognitionMeihuizi Jia, Xin Shen, Lei Shen, Jinhui Pang et al.ACM MM 2022 · 45 citations
- MCG-MNER: A Multi-Granularity Cross-Modality Generative Framework for Multimodal NER with InstructionJunjie Wu, Chen Gong, Ziqiang Cao, Guohong FuACM MM 2023 · 14 citations
- Ground and Reconstruct: Entity-Region Bidirectional Alignment Pre-Training for Low-Resource GMNERRunwei Situ, Yi Cai, Yong Xu, Jiexin WangACM MM 2025 · 3 citations
