A Novel Energy Based Model Mechanism for Multi-Modal Aspect-Based Sentiment Analysis
Tianshuo Peng, Zuchao Li, Ping Wang, Lefei Zhang, Hai Zhao
Abstract
Multi-modal aspect-based sentiment analysis (MABSA) has recently attracted increasing attention. The span-based extraction methods, such as FSUIE, demonstrate strong performance in sentiment analysis due to their joint modeling of input sequences and target labels. However, previous methods still have certain limitations: (i) They ignore the difference in the focus of visual information between different analysis targets (aspect or sentiment). (ii) Combining features from uni-modal encoders directly may not be sufficient to eliminate the modal gap and can cause difficulties in capturing the image-text pairwise relevance. (iii) Existing span-based methods for MABSA ignore the pairwise relevance of target span boundaries. To tackle these limitations, we propose a novel framework called DQPSA for multi-modal sentiment analysis. Specifically, our model contains a Prompt as Dual Query (PDQ) module that uses the prompt as both a visual query and a language query to extract prompt-aware visual information and strengthen the pairwise relevance between visual information and the analysis target. Additionally, we introduce an Energy-based Pairwise Expert (EPE) module that models the boundaries pairing of the analysis target from the perspective of an Energy-based Model. This expert predicts aspect or sentiment span based on pairwise stability. Experiments on three widely used benchmarks demonstrate that DQPSA outperforms previous approaches and achieves a new state-of-the-art performance. Furthermore, we conducted a fair comparison with relevant large-scale models such as ChatGPT-3.5 and VisualGLM. We found that our model has * Corresponding author. † Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- DEQA: Descriptions Enhanced Question-Answering Framework for Multimodal Aspect-Based Sentiment AnalysisZhixin Han, Mengting Hu, Yinhao Bai, Xunzhi Wang et al.AAAI 2025 · 3 citations
- An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception CapabilityDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma et al.ICML 2025
- OX-MABSR: A Benchmark for Open-domain Explainable Multimodal Aspect-Based Sentiment ReasoningXinjing Liu, Zixin Xue, Pengyue Lin, Xinyu Tu et al.AAAI 2026
- Deep Incomplete Multi-View Clustering via Hierarchical Imputation and AlignmentYiming Du, Ziyu Wang, Jian Li, Rui Ning et al.AAAI 2026
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Position-Aware Tagging for Aspect Sentiment Triplet ExtractionLu Xu, Hao Li, Wei Lu, Lidong BingEMNLP 2020 · 264 citations
- Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerJianfei Yu, Jing Jiang, Li Yang, Rui XiaACL 2020 · 260 citations
- Exploiting BERT for Multimodal Target Sentiment Classification through Input Space TranslationZaid Khan, Yun FuACM MM 2021 · 192 citations
Related papers
- Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment AnalysisYan Ling, Jianfei Yu, Rui XiaACL 2022 · 116 citations
- Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation DetectionXincheng Ju, Dong Zhang, Rong Xiao, Junhui Li et al.EMNLP 2021 · 130 citations
- Aspect Enhancement and Text Simplification in Multimodal Aspect-Based Sentiment Analysis for Multi-Aspect and Multi-Sentiment ScenariosLinlin Zhu, Heli Sun, Qunshu Gao, Yuze Liu et al.AAAI 2025 · 10 citations
- Aspects are Anchors: Towards Multimodal Aspect-based Sentiment Analysis via Aspect-driven Alignment and RefinementZhanpeng Chen, Zhihong Zhu, Wanshi Xu, Yunyan Zhang et al.ACM MM 2024 · 8 citations
- Span-Pair Interaction and Tagging for Dialogue-Level Aspect-Based Sentiment Quadruple AnalysisChangzhi Zhou, Zhijing Wu, Dandan Song, Linmei Hu et al.WWW 2024 · 8 citations
