Fine-Grained Multimodal Named Entity Recognition and Grounding with a Generative Framework
Jieming Wang, Ziyan Li, Jianfei Yu, Li Yang, Rui Xia
摘要
Multimodal Named Entity Recognition (MNER) aims to locate and classify named entities mentioned in a pair of text and image. However, most previous MNER works focus on extracting entities in the form of text but failing to ground text symbols to their corresponding visual objects. Moreover, existing MNER studies primarily classify entities into four coarse-grained entity types, which are often insufficient to map them to their real-world referents. To solve these limitations, we introduce a task named Fine-grained Multimodal Named Entity Recognition and Grounding (FMNERG) in this paper, which aims to simultaneously extract named entities in text, their fine-grained entity types, and their grounded visual objects in image. Moreover, we construct a Twitter dataset for the FMNERG task, and further propose a T5-based multImodal GEneration fRamework (TIGER), which formulates FMNERG as a generation problem by converting all the entity-type-object triples into a target sequence and adapts a pre-trained sequence-to-sequence model T5 to directly generate the target sequence from an image-text input pair. Experimental results demonstrate that TIGER performs significantly better than a number of baseline systems on the annotated Twitter dataset. Our dataset annotation and source code are publicly released at https://github.com/NUSTM/FMNERG.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Cross-modal Multi-task Learning for Multimedia Event ExtractionJianwei Cao, Yanli Hu, Zhen Tan, Xiang ZhaoAAAI 2025 · 被引用 8 次
- Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity RecognitionJielong Tang, Zhenxing Wang, Ziyang Gong, Jianxing Yu 等AAAI 2025 · 被引用 8 次
- SAKE: Self-aware Knowledge Exploitation-Exploration for Grounded Multimodal Named Entity RecognitionJielong Tang, Xujie Yuan, Jiayang Liu, Jianxing Yu 等KDD 2026 · 被引用 1 次
- ISR: Self-Refining Referring Expressions for Entity GroundingZhuocheng Yu, Bingchan Zhao, Yifan Song, Sujian Li 等ACL 2025 · 被引用 1 次
- MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity RecognitionXinkui Lin, Yuhui Zhang, Yongxiu Xu, Kun Huang 等EMNLP 2025
相关 Paper
- Grounded Multimodal Named Entity Recognition on Social MediaJianfei Yu, Ziyan Li, Jieming Wang, Rui XiaACL 2023 · 被引用 32 次
- MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query GroundingMeihuizi Jia, Lei Shen, Xin Shen, Lejian Liao 等AAAI 2023 · 被引用 68 次
- Query Prior Matters: A MRC Framework for Multimodal Named Entity RecognitionMeihuizi Jia, Xin Shen, Lei Shen, Jinhui Pang 等ACM MM 2022 · 被引用 45 次
- MCG-MNER: A Multi-Granularity Cross-Modality Generative Framework for Multimodal NER with InstructionJunjie Wu, Chen Gong, Ziqiang Cao, Guohong FuACM MM 2023 · 被引用 14 次
- Ground and Reconstruct: Entity-Region Bidirectional Alignment Pre-Training for Low-Resource GMNERRunwei Situ, Yi Cai, Yong Xu, Jiexin WangACM MM 2025 · 被引用 3 次
