Ground and Reconstruct: Entity-Region Bidirectional Alignment Pre-Training for Low-Resource GMNER
Runwei Situ, Yi Cai, Yong Xu, Jiexin Wang
摘要
Grounded Multimodal Named Entity Recognition (GMNER) extends Multimodal Named Entity Recognition (MNER) by identifying named entities, their types, and corresponding image regions. Fine-grained MNER and Grounding (FMNERG) further refines entity categorization. However, existing methods struggle with scarce annotated data, particularly in low-resource scenarios, and often fail to generalize to unseen entities. While vision-language pre-training (VLP) leverages unlabeled image-caption pairs, it primarily learns generic visual-linguistic representations, overlooking fine-grained entity-region alignment crucial for entity-related tasks. To address these challenges, we propose a unified VLP framework for GMNER and FMNERG, introducing two task-specific pre-training objectives: Entity-to-Region Alignment (ETRA) for entity grounding and Region-to-Entity Alignment (RTEA) for entity reconstruction. These tasks jointly optimize fine-grained entity-region alignment. To compensate for the lack of fine-grained multimodal pre-training data, we develop an automatic labeling method that distills entity-oriented knowledge from large-scale unlabeled image-text pairs, enhancing generalization to unseen entities. Extensive experiments on GMNER and FMNERG benchmarks demonstrate that our framework outperforms existing low-resource learning approaches and achieves competitive performance in full-supervision, underscoring its effectiveness across diverse data conditions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Fine-Grained Multimodal Named Entity Recognition and Grounding with a Generative FrameworkJieming Wang, Ziyan Li, Jianfei Yu, Li Yang 等ACM MM 2023 · 被引用 11 次
- UnCo: Uncertainty-Driven Collaborative Framework of Large and Small Models for Grounded Multimodal NERJielong Tang, Yang Yang, Jianxing Yu, Zhen-Xing Wang 等EMNLP 2025 · 被引用 3 次
- Grounded Multimodal Named Entity Recognition on Social MediaJianfei Yu, Ziyan Li, Jieming Wang, Rui XiaACL 2023 · 被引用 32 次
- Unsupervised Vision-and-Language Pretraining via Retrieval-based Multi-Granular AlignmentMingyang Zhou, Licheng Yu, Amanpreet Singh, Mengjiao Wang 等CVPR 2022 · 被引用 29 次
- MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query GroundingMeihuizi Jia, Lei Shen, Xin Shen, Lejian Liao 等AAAI 2023 · 被引用 68 次
