Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis
Yan Ling, Jianfei Yu, Rui Xia
Abstract
As an important task in sentiment analysis, Multimodal Aspect-Based Sentiment Analysis (MABSA) has attracted increasing attention in recent years. However, previous approaches either (i) use separately pre-trained visual and textual models, which ignore the crossmodal alignment or (ii) use vision-language models pre-trained with general pre-training tasks, which are inadequate to identify finegrained aspects, opinions, and their alignments across modalities. To tackle these limitations, we propose a task-specific Vision-Language Pre-training framework for MABSA (VLP-MABSA), which is a unified multimodal encoder-decoder architecture for all the pretraining and downstream tasks. We further design three types of task-specific pre-training tasks from the language, vision, and multimodal modalities, respectively. Experimental results show that our approach generally outperforms the state-of-the-art approaches on three MABSA subtasks. Further analysis demonstrates the effectiveness of each pretraining task. The source code is publicly released at https://github.com/NUSTM/ VLP-MABSA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7bb4688-96e1-4bd1-b71a-7a82c5e2b7b0Cited by top-tier papers19
- All in One: Multi-Task Prompting for Graph Neural NetworksXiangguo Sun, Hong Cheng, Jia Li, Bo Liu et al.KDD 2023 · 149 citations
- Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment AnalysisHao Yang, Yanyan Zhao, Bing QinEMNLP 2022 · 60 citations
- Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionDongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu OkumuraEMNLP 2023 · 46 citations
- M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment AnalysisFei Zhao, Chunhui Li, Zhen Wu, Yawen Ouyang et al.EMNLP 2023 · 42 citations
- WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World KnowledgeWenbin Wang, Liang Ding, Li Shen, Yong Luo et al.ACM MM 2024 · 39 citations
Builds on11
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun et al.AAAI 2021 · 414 citations
- Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerJianfei Yu, Jing Jiang, Li Yang, Rui XiaACL 2020 · 260 citations
- Exploiting BERT for Multimodal Target Sentiment Classification through Input Space TranslationZaid Khan, Yun FuACM MM 2021 · 192 citations
- RpBERT: A Text-image Relation Propagation-based BERT Model for Multimodal NERLin Sun, Jiquan Wang, Kai Zhang, Yindu Su et al.AAAI 2021 · 189 citations
- Multimodal Representation with Embedded Visual Guiding Objects for Named Entity Recognition in Social Media PostsZhiwei Wu, Changmeng Zheng, Yi Cai, Junying Chen et al.ACM MM 2020 · 139 citations
Related papers
- UniSA: Unified Generative Framework for Sentiment AnalysisZaijing Li, Ting-En Lin, Yuchuan Wu, Meng Liu et al.ACM MM 2023 · 22 citations
- A Novel Energy Based Model Mechanism for Multi-Modal Aspect-Based Sentiment AnalysisTianshuo Peng, Zuchao Li, Ping Wang, Lefei Zhang et al.AAAI 2024 · 22 citations
- Probing Sentiment-Oriented PreTraining Inspired by Human Sentiment Perception MechanismTinglei Feng, Jiaxuan Liu, Jufeng YangCVPR 2023
- Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation DetectionXincheng Ju, Dong Zhang, Rong Xiao, Junhui Li et al.EMNLP 2021 · 130 citations
- Unified Multi-modal Pre-training for Few-shot Sentiment Analysis with Prompt-based LearningYang Yu, Dong Zhang, Shoushan LiACM MM 2022 · 44 citations
