Lune

ACM MM2025Top-tier venue

Video-Level Multimodal Relation Extraction with Event-Entity Semantic Consistency

Zefan Zhang, Weiqi Zhang, Kailong Suo, Yanhui Li, Tian Bai

2025Year

Abstract

Previous research on Multimodal Relation Extraction (MRE) has primarily focused on identifying textual relations enhanced by static visual clues from images, benefiting fields such as multimedia analysis and knowledge graphs. With the rapid rise of video content on social media platforms, Multimodal Relation Extraction (MRE) systems face new challenges. To bridge this gap, we introduce Video-level Multimodal Relation Extraction (VMRE), a novel task aimed at extracting relational facts from videos. To advance this research, we present Vid-MRE, a new dataset containing 32 relation types and 12,402 multimodal relational facts, annotated across 3,970 pairs of textual news titles and corresponding videos. Since this task demands precise event and entity grounding to filter out excessive noise in the video, we propose an Event-Entity Semantic Consistency Network (E2SCN) to capture relational clues in the video effectively. Experimental results demonstrate that incorporating video content into the model significantly improves relation identification performance but also introduces more noise. Our E2SCN method effectively reduces the noise, enhancing fine-grained multimodal event and entity alignments while achieving state-of-the-art (SOTA) performance.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get af0b137e-bda2-44f4-b606-61b4f809d44f

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines