Lune

ACM MM2025顶会

Ground and Reconstruct: Entity-Region Bidirectional Alignment Pre-Training for Low-Resource GMNER

Runwei Situ, Yi Cai, Yong Xu, Jiexin Wang

2025年份
3被引次数

摘要

Grounded Multimodal Named Entity Recognition (GMNER) extends Multimodal Named Entity Recognition (MNER) by identifying named entities, their types, and corresponding image regions. Fine-grained MNER and Grounding (FMNERG) further refines entity categorization. However, existing methods struggle with scarce annotated data, particularly in low-resource scenarios, and often fail to generalize to unseen entities. While vision-language pre-training (VLP) leverages unlabeled image-caption pairs, it primarily learns generic visual-linguistic representations, overlooking fine-grained entity-region alignment crucial for entity-related tasks. To address these challenges, we propose a unified VLP framework for GMNER and FMNERG, introducing two task-specific pre-training objectives: Entity-to-Region Alignment (ETRA) for entity grounding and Region-to-Entity Alignment (RTEA) for entity reconstruction. These tasks jointly optimize fine-grained entity-region alignment. To compensate for the lack of fine-grained multimodal pre-training data, we develop an automatic labeling method that distills entity-oriented knowledge from large-scale unlabeled image-text pairs, enhancing generalization to unseen entities. Extensive experiments on GMNER and FMNERG benchmarks demonstrate that our framework outperforms existing low-resource learning approaches and achieves competitive performance in full-supervision, underscoring its effectiveness across diverse data conditions.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖