Video-Level Multimodal Relation Extraction with Event-Entity Semantic Consistency
Zefan Zhang, Weiqi Zhang, Kailong Suo, Yanhui Li, Tian Bai
Abstract
Previous research on Multimodal Relation Extraction (MRE) has primarily focused on identifying textual relations enhanced by static visual clues from images, benefiting fields such as multimedia analysis and knowledge graphs. With the rapid rise of video content on social media platforms, Multimodal Relation Extraction (MRE) systems face new challenges. To bridge this gap, we introduce Video-level Multimodal Relation Extraction (VMRE), a novel task aimed at extracting relational facts from videos. To advance this research, we present Vid-MRE, a new dataset containing 32 relation types and 12,402 multimodal relational facts, annotated across 3,970 pairs of textual news titles and corresponding videos. Since this task demands precise event and entity grounding to filter out excessive noise in the video, we propose an Event-Entity Semantic Consistency Network (E2SCN) to capture relational clues in the video effectively. Experimental results demonstrate that incorporating video content into the model significantly improves relation identification performance but also introduces more noise. Our E2SCN method effectively reduces the noise, enhancing fine-grained multimodal event and entity alignments while achieving state-of-the-art (SOTA) performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get af0b137e-bda2-44f4-b606-61b4f809d44fRelated papers
- Multimodal Relation Extraction with Efficient Graph AlignmentChangmeng Zheng, Junhao Feng, Ze Fu, Yi Cai et al.ACM MM 2021 · 134 citations
- MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark EvaluationLiang He, Hongke Wang, Yongchang Cao, Zhen Wu et al.ACM MM 2023 · 17 citations
- A Hierarchical Network for Multimodal Document-Level Relation ExtractionLingxing Kong, Jiuliang Wang, Zheng Ma, Qifeng Zhou et al.AAAI 2024 · 3 citations
- Joint Multimodal Entity-Relation Extraction Based on Edge-Enhanced Graph Alignment Network and Word-Pair Relation TaggingLi Yuan, Yi Cai, Jin Wang, Qing LiAAAI 2023 · 89 citations
- REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-ExpertsXinkui Lin, Yongxiu Xu, Minghao Tang, Shilong Zhang et al.ACM MM 2025 · 2 citations
