A Hierarchical Network for Multimodal Document-Level Relation Extraction
Lingxing Kong, Jiuliang Wang, Zheng Ma, Qifeng Zhou, Jianbing Zhang, Liang He, Jiajun Chen
摘要
Document-level relation extraction aims to extract entity relations that span across multiple sentences. This task faces two critical issues: long dependency and mention selection. Prior works address the above problems from the textual perspective, however, it is hard to handle these problems solely based on text information. In this paper, we leverage video information to provide additional evidence for understanding long dependencies and offer a wider perspective for identifying relevant mentions, thus giving rise to a new task named Multimodal Document-level Relation Extraction (MDocRE). To tackle this new task, we construct a human-annotated dataset including documents and relevant videos, which, to the best of our knowledge, is the first document-level relation extraction dataset equipped with video clips. We also propose a hierarchical framework to learn interactions between different dependency levels and a textual-guided transformer architecture that incorporates both textual and video modalities. In addition, we utilize a mention gate module to address the mention-selection problem in both modalities. Experiments on our proposed dataset show that 1) incorporating video information greatly improves model performance; 2) our hierarchical framework has state-of-the-art results compared with both unimodal and multimodal baselines; 3) through collaborating with video information, our model better solves the long-dependency and mention-selection problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Reasoning with Latent Structure Refinement for Document-Level Relation ExtractionGuoshun Nan, Zhijiang Guo, Ivan Sekulic, Wei LuACL 2020 · 被引用 294 次
- Double Graph Based Reasoning for Document-level Relation ExtractionShuang Zeng, Runxin Xu, Baobao Chang, Lei LiEMNLP 2020 · 被引用 238 次
- Entity Structure Within and Throughout: Modeling Mention Dependencies for Document-Level Relation ExtractionBenfeng Xu, Quan Wang, Yajuan Lyu, Yong Zhu 等AAAI 2021 · 被引用 200 次
- Global-to-Local Neural Networks for Document-Level Relation ExtractionDifeng Wang, Wei Hu, Ermei Cao, Weijian SunEMNLP 2020 · 被引用 122 次
相关 Paper
- Revisiting Document-Level Relation Extraction with Context-Guided Link PredictionMonika Jain, Raghava Mutharaju, Ramakanth Kavuluru, Kuldeep SinghAAAI 2024 · 被引用 17 次
- Video-Level Multimodal Relation Extraction with Event-Entity Semantic ConsistencyZefan Zhang, Weiqi Zhang, Kailong Suo, Yanhui Li 等ACM MM 2025
- Anaphor Assisted Document-Level Relation ExtractionChonggang Lu, Richong Zhang, Kai Sun, Jaein Kim 等EMNLP 2023 · 被引用 16 次
- MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark EvaluationLiang He, Hongke Wang, Yongchang Cao, Zhen Wu 等ACM MM 2023 · 被引用 17 次
- REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-ExpertsXinkui Lin, Yongxiu Xu, Minghao Tang, Shilong Zhang 等ACM MM 2025 · 被引用 2 次
