RpBERT: A Text-image Relation Propagation-based BERT Model for Multimodal NER
Lin Sun, Jiquan Wang, Kai Zhang, Yindu Su, Fangsheng Weng
摘要
Recently multimodal named entity recognition (MNER) has utilized images to improve the accuracy of NER in tweets. However, most of the multimodal methods use attention mechanisms to extract visual clues regardless of whether the text and image are relevant. Practically, the irrelevant textimage pairs account for a large proportion in tweets. The visual clues that are unrelated to the texts will exert uncertain or even negative effects on multimodal model learning. In this paper, we introduce a method of text-image relation propagation into the multimodal BERT model. We integrate soft or hard gates to select visual clues and propose a multitask algorithm to train on the MNER datasets. In the experiments, we deeply analyze the changes in visual attention before and after the use of text-image relation propagation. Our model achieves state-of-the-art performance on the MNER datasets. The source code is available online 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng 等SIGIR 2022 · 被引用 227 次
- Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment AnalysisYan Ling, Jianfei Yu, Rui XiaACL 2022 · 被引用 116 次
- Learning from Different text-image Pairs: A Relation-enhanced Graph Convolutional Network for Multimodal NERFei Zhao, Chunhui Li, Zhen Wu, Shangyu Xing 等ACM MM 2022 · 被引用 59 次
- M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment AnalysisFei Zhao, Chunhui Li, Zhen Wu, Yawen Ouyang 等EMNLP 2023 · 被引用 42 次
- Grounded Multimodal Named Entity Recognition on Social MediaJianfei Yu, Ziyan Li, Jieming Wang, Rui XiaACL 2023 · 被引用 32 次
它引用的顶会 Paper3
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerJianfei Yu, Jing Jiang, Li Yang, Rui XiaACL 2020 · 被引用 260 次
相关 Paper
- MCG-MNER: A Multi-Granularity Cross-Modality Generative Framework for Multimodal NER with InstructionJunjie Wu, Chen Gong, Ziqiang Cao, Guohong FuACM MM 2023 · 被引用 14 次
- MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query GroundingMeihuizi Jia, Lei Shen, Xin Shen, Lejian Liao 等AAAI 2023 · 被引用 68 次
- Hierarchical Aligned Multimodal Learning for NER on Tweet PostsPeipei Liu, Hong Li, Yimo Ren, Jie Liu 等AAAI 2024 · 被引用 11 次
- Query Prior Matters: A MRC Framework for Multimodal Named Entity RecognitionMeihuizi Jia, Xin Shen, Lei Shen, Jinhui Pang 等ACM MM 2022 · 被引用 45 次
- Learning Implicit Entity-object Relations by Bidirectional Generative Alignment for Multimodal NERFeng Chen, Jiajia Liu, Kaixiang Ji, Wang Ren 等ACM MM 2023 · 被引用 14 次
