Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis
Yan Ling, Jianfei Yu, Rui Xia
摘要
As an important task in sentiment analysis, Multimodal Aspect-Based Sentiment Analysis (MABSA) has attracted increasing attention in recent years. However, previous approaches either (i) use separately pre-trained visual and textual models, which ignore the crossmodal alignment or (ii) use vision-language models pre-trained with general pre-training tasks, which are inadequate to identify finegrained aspects, opinions, and their alignments across modalities. To tackle these limitations, we propose a task-specific Vision-Language Pre-training framework for MABSA (VLP-MABSA), which is a unified multimodal encoder-decoder architecture for all the pretraining and downstream tasks. We further design three types of task-specific pre-training tasks from the language, vision, and multimodal modalities, respectively. Experimental results show that our approach generally outperforms the state-of-the-art approaches on three MABSA subtasks. Further analysis demonstrates the effectiveness of each pretraining task. The source code is publicly released at https://github.com/NUSTM/ VLP-MABSA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- All in One: Multi-Task Prompting for Graph Neural NetworksXiangguo Sun, Hong Cheng, Jia Li, Bo Liu 等KDD 2023 · 被引用 149 次
- Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment AnalysisHao Yang, Yanyan Zhao, Bing QinEMNLP 2022 · 被引用 60 次
- Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionDongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu OkumuraEMNLP 2023 · 被引用 46 次
- M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment AnalysisFei Zhao, Chunhui Li, Zhen Wu, Yawen Ouyang 等EMNLP 2023 · 被引用 42 次
- WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World KnowledgeWenbin Wang, Liang Ding, Li Shen, Yong Luo 等ACM MM 2024 · 被引用 39 次
它引用的顶会 Paper11
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun 等AAAI 2021 · 被引用 414 次
- Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerJianfei Yu, Jing Jiang, Li Yang, Rui XiaACL 2020 · 被引用 260 次
- Exploiting BERT for Multimodal Target Sentiment Classification through Input Space TranslationZaid Khan, Yun FuACM MM 2021 · 被引用 192 次
- RpBERT: A Text-image Relation Propagation-based BERT Model for Multimodal NERLin Sun, Jiquan Wang, Kai Zhang, Yindu Su 等AAAI 2021 · 被引用 189 次
- Multimodal Representation with Embedded Visual Guiding Objects for Named Entity Recognition in Social Media PostsZhiwei Wu, Changmeng Zheng, Yi Cai, Junying Chen 等ACM MM 2020 · 被引用 139 次
相关 Paper
- UniSA: Unified Generative Framework for Sentiment AnalysisZaijing Li, Ting-En Lin, Yuchuan Wu, Meng Liu 等ACM MM 2023 · 被引用 22 次
- A Novel Energy Based Model Mechanism for Multi-Modal Aspect-Based Sentiment AnalysisTianshuo Peng, Zuchao Li, Ping Wang, Lefei Zhang 等AAAI 2024 · 被引用 22 次
- Probing Sentiment-Oriented PreTraining Inspired by Human Sentiment Perception MechanismTinglei Feng, Jiaxuan Liu, Jufeng YangCVPR 2023
- Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation DetectionXincheng Ju, Dong Zhang, Rong Xiao, Junhui Li 等EMNLP 2021 · 被引用 130 次
- Unified Multi-modal Pre-training for Few-shot Sentiment Analysis with Prompt-based LearningYang Yu, Dong Zhang, Shoushan LiACM MM 2022 · 被引用 44 次
