U-MERE: Unconstrained Multimodal Entity and Relation Extraction with Collaborative Modeling and Order-Sensitive Optimization
Wei Jia, Li Jin, Kaiwen Wei, Yuying Shang, Nayu Liu, Zhicong Lu, Qing Liu, Linhao Zhang, Jiang Zhong, Yanfeng Hu
Abstract
Existing multimodal entity and relation extraction tasks primarily focus on text-to-text or text-to-visual entity relations, overlooking real-world complexities involving visual-to-text and visual-to-visual cases, thus failing to capture the richer semantic structures in complex cross-modal interactions. To address the limitations, we propose a new task, Unconstrained Multimodal Entity and Relation Extraction (U-MERE), which jointly extracts arbitrary visual and textual entities, and their relations from image-text pairs. To accomplish U-MERE, we construct UMERE-Bench, a benchmark with over 9,000 samples that comprehensively covers four cross-modal entity relation directions and three task settings. Given the difficulty of jointly modeling diverse directions of cross-modal entity relations, we introduce Collaborative Modeling and Order-Sensitive (CMOS), which collaboratively guides large vision-language models (LVLMs) to decompose task complexity and mitigates generation order bias from fixed target relation sequences. CMOS employs small models to generate candidate entities, guiding LVLMs to capture key information and jointly optimizes multiple feasible relation orderings to reduce order dependency. Additionally, we design a Multimodal Order-aware Matching (MOM) evaluation method to align predictions with ground truth for precise assessment. Experimental results reveal that current LVLMs show limited performance on U-MERE, underscoring its inherent challenges, while CMOS consistently achieves superior performance across multiple advanced LVLMs, demonstrating its effectiveness and generalization capability. The dataset and code will be available in https://github.com/jiaweidoris/U-MERE.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3df7e577-1ec2-45f2-bccd-0758dd586c89Cited by top-tier papers1
Ask how each one uses itRelated papers
- Caption-Aware Multimodal Relation Extraction with Mutual Information MaximizationZefan Zhang, Weiqi Zhang, Yanhui Li, Tian BaiACM MM 2024 · 9 citations
- CoMCo: Consistency-Aware Multi-Agent Coordination for Zero-Shot Cross-Modal Entity MatchingShiqi Zhang, Weixin Zeng, Ziheng Zhang, Wenzhe Hou et al.SIGIR 2026
- MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark EvaluationLiang He, Hongke Wang, Yongchang Cao, Zhen Wu et al.ACM MM 2023 · 17 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity LinkingZiyan Liu, Junwen Li, Kaiwen Li, Tong Ruan et al.ACM MM 2025 · 2 citations
