Lune

ACM MM2025Top-tier venue

Ground and Reconstruct: Entity-Region Bidirectional Alignment Pre-Training for Low-Resource GMNER

Runwei Situ, Yi Cai, Yong Xu, Jiexin Wang

2025Year
3Citations

Abstract

Grounded Multimodal Named Entity Recognition (GMNER) extends Multimodal Named Entity Recognition (MNER) by identifying named entities, their types, and corresponding image regions. Fine-grained MNER and Grounding (FMNERG) further refines entity categorization. However, existing methods struggle with scarce annotated data, particularly in low-resource scenarios, and often fail to generalize to unseen entities. While vision-language pre-training (VLP) leverages unlabeled image-caption pairs, it primarily learns generic visual-linguistic representations, overlooking fine-grained entity-region alignment crucial for entity-related tasks. To address these challenges, we propose a unified VLP framework for GMNER and FMNERG, introducing two task-specific pre-training objectives: Entity-to-Region Alignment (ETRA) for entity grounding and Region-to-Entity Alignment (RTEA) for entity reconstruction. These tasks jointly optimize fine-grained entity-region alignment. To compensate for the lack of fine-grained multimodal pre-training data, we develop an automatic labeling method that distills entity-oriented knowledge from large-scale unlabeled image-text pairs, enhancing generalization to unseen entities. Extensive experiments on GMNER and FMNERG benchmarks demonstrate that our framework outperforms existing low-resource learning approaches and achieves competitive performance in full-supervision, underscoring its effectiveness across diverse data conditions.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines