A Hierarchical Network for Multimodal Document-Level Relation Extraction
Lingxing Kong, Jiuliang Wang, Zheng Ma, Qifeng Zhou, Jianbing Zhang, Liang He, Jiajun Chen
Abstract
Document-level relation extraction aims to extract entity relations that span across multiple sentences. This task faces two critical issues: long dependency and mention selection. Prior works address the above problems from the textual perspective, however, it is hard to handle these problems solely based on text information. In this paper, we leverage video information to provide additional evidence for understanding long dependencies and offer a wider perspective for identifying relevant mentions, thus giving rise to a new task named Multimodal Document-level Relation Extraction (MDocRE). To tackle this new task, we construct a human-annotated dataset including documents and relevant videos, which, to the best of our knowledge, is the first document-level relation extraction dataset equipped with video clips. We also propose a hierarchical framework to learn interactions between different dependency levels and a textual-guided transformer architecture that incorporates both textual and video modalities. In addition, we utilize a mention gate module to address the mention-selection problem in both modalities. Experiments on our proposed dataset show that 1) incorporating video information greatly improves model performance; 2) our hierarchical framework has state-of-the-art results compared with both unimodal and multimodal baselines; 3) through collaborating with video information, our model better solves the long-dependency and mention-selection problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 860c8497-7e8e-4eb7-9baa-06ea5830ead8Builds on7
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Reasoning with Latent Structure Refinement for Document-Level Relation ExtractionGuoshun Nan, Zhijiang Guo, Ivan Sekulic, Wei LuACL 2020 · 294 citations
- Double Graph Based Reasoning for Document-level Relation ExtractionShuang Zeng, Runxin Xu, Baobao Chang, Lei LiEMNLP 2020 · 238 citations
- Entity Structure Within and Throughout: Modeling Mention Dependencies for Document-Level Relation ExtractionBenfeng Xu, Quan Wang, Yajuan Lyu, Yong Zhu et al.AAAI 2021 · 200 citations
- Global-to-Local Neural Networks for Document-Level Relation ExtractionDifeng Wang, Wei Hu, Ermei Cao, Weijian SunEMNLP 2020 · 122 citations
Related papers
- Revisiting Document-Level Relation Extraction with Context-Guided Link PredictionMonika Jain, Raghava Mutharaju, Ramakanth Kavuluru, Kuldeep SinghAAAI 2024 · 17 citations
- Video-Level Multimodal Relation Extraction with Event-Entity Semantic ConsistencyZefan Zhang, Weiqi Zhang, Kailong Suo, Yanhui Li et al.ACM MM 2025
- Anaphor Assisted Document-Level Relation ExtractionChonggang Lu, Richong Zhang, Kai Sun, Jaein Kim et al.EMNLP 2023 · 16 citations
- MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark EvaluationLiang He, Hongke Wang, Yongchang Cao, Zhen Wu et al.ACM MM 2023 · 17 citations
- REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-ExpertsXinkui Lin, Yongxiu Xu, Minghao Tang, Shilong Zhang et al.ACM MM 2025 · 2 citations
