Multimodal Relation Extraction via a Mixture of Hierarchical Visual Context Learners
Xiyang Liu, Chunming Hu, Richong Zhang, Kai Sun, Samuel Mensah, Yongyi Mao
摘要
Multimodal relation extraction is a fundamental task of multimodal information extraction. Recent studies have shown promising results by integrating hierarchical visual features from local regions, like image patches, to the broader global regions that form the entire image. However, research to date has largely ignored the understanding of how hierarchical visual semantics are represented and the characteristics that can benefit relation extraction. To bridge this gap, we propose a novel two-stage hierarchical visual context fusion transformer incorporating the mixture of multimodal experts framework to effectively represent and integrate hierarchical visual features into textual semantic representations. In addition, we introduce the concept of hierarchical tracking maps to facilitate the understanding of the intrinsic mechanisms of image information processing involved in multimodal models. We thoroughly investigate the implications of hierarchical visual contexts through four dimensions: performance evaluation, the nature of auxiliary visual information, the patterns observed in the image encoding hierarchy, and the significance of various visual encoding levels. Empirical studies show that our approach achieves new state-ofthe-art performance on the MNRE dataset. 1 CCS CONCEPTS • Information systems → Multimedia and multimodal retrieval; • Computing methodologies → Information extraction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Graph Mixture of Experts and Memory-augmented Routers for Multivariate Time Series Anomaly DetectionXiaoyu Huang, Weidong Chen, Bo Hu, Zhendong MaoAAAI 2025 · 被引用 22 次
- Prototype-Guided Multimodal Relation Extraction based on Entity AttributesZefan Zhang, Weiqi Zhang, Yanhui Li, Tian BaiAAAI 2025 · 被引用 8 次
- REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-ExpertsXinkui Lin, Yongxiu Xu, Minghao Tang, Shilong Zhang 等ACM MM 2025 · 被引用 2 次
- Retrieval over Classification: Integrating Relation Semantics for Multimodal Relation ExtractionLei Hei, Tingjing Liao, Peiyingxin, Yiyang Qi 等EMNLP 2025
- VisionST: Coordinating Cross-modal Traffic Prediction with Interactive Geo-image EncodingJinwen Chen, Hao Miao, Chenxi Liu, Yan Zhao 等WWW 2026
它引用的顶会 Paper17
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang 等ICCV 2019 · 被引用 441 次
相关 Paper
- A Hierarchical Network for Multimodal Document-Level Relation ExtractionLingxing Kong, Jiuliang Wang, Zheng Ma, Qifeng Zhou 等AAAI 2024 · 被引用 3 次
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng 等SIGIR 2022 · 被引用 227 次
- MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion RecognitionJian Chen, Yuxuan Hu, Haifeng Lu, Wei Wang 等ACM MM 2025 · 被引用 5 次
- Multi-Information Hierarchical Fusion Transformer with Local Alignment and Global Correlation for Micro-Expression RecognitionJinsheng Wei, Jialiang Sun, Guanming Lu, Jingjie Yan 等ACM MM 2025 · 被引用 5 次
- Multimodal High-order Relation Transformer for Scene Boundary DetectionXi Wei, Zhangxiang Shi, Tianzhu Zhang, Xiaoyuan Yu 等ICCV 2023 · 被引用 7 次
