Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cues
Wei Chen, Tongguan Wang, Feiyue Xue, Junkai Li, Hui Liu, Ying Sha
摘要
Multimodal desire understanding, a task closely related to both emotion and sentiment that aims to infer human intentions from visual and textual cues, is an emerging yet underexplored task in affective computing with applications in social media analysis. Existing methods for related tasks predominantly focus on mining verbal cues, often overlooking the effective utilization of non-verbal cues embedded in images. To bridge this gap, we propose a Symmetrical Bidirectional Multimodal Learning Framework for Desire, Emotion, and Sentiment Recognition (SyDES). The core of SyDES is to achieve bidirectional fine-grained modal alignment between text and image modalities. Specifically, we introduce a mixed-scaled image strategy that combines global context from low-resolution images with fine-grained local features via masked image modeling (MIM) on high-resolution sub-images, effectively capturing intention-related visual representations. Then, we devise symmetrical cross-modal decoders, including a text-guided image decoder and an image-guided text decoder, which enable mutual reconstruction and refinement between modalities, facilitating deep cross-modal interaction. Furthermore, a set of dedicated loss functions is designed to harmonize potential conflicts between the MIM and modal alignment objectives during optimization. Extensive evaluations on the MSED benchmark demonstrate the superiority of our approach, which establishes a new state-of-the-art performance with 1.1% F1-score improvement in desire understanding. Consistent gains in emotion and sentiment recognition further validate its generalization ability and the necessity of utilizing non-verbal cues. Our code is available at: https://github.com/especiallyW/SyDES .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World KnowledgeWenbin Wang, Liang Ding, Li Shen, Yong Luo 等ACM MM 2024 · 被引用 39 次
- Multimodal Sentiment Detection Based on Multi-channel Graph Neural NetworksXiaocui Yang, Shi Feng, Yifei Zhang, Daling WangACL 2021
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li 等CVPR 2022
相关 Paper
- Generating Multimodal Metaphorical Features for Meme UnderstandingBo Xu, Junzhe Zheng, Jiayuan He, Yuxuan Sun 等ACM MM 2024 · 被引用 6 次
- Multi-Granular Multimodal Clue Fusion for Meme UnderstandingLi Zheng, Hao Fei, Ting Dai, Zuquan Peng 等AAAI 2025 · 被引用 13 次
- Tri-Subspaces Disentanglement for Multimodal Sentiment AnalysisChunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu 等CVPR 2026 · 被引用 7 次
- Uncertain Multimodal Intention and Emotion Understanding in the WildQu Yang, Qinghongya Shi, Tongxin Wang, Mang YeCVPR 2025
- Look, Read and Feel: Benchmarking Ads Understanding with Multimodal Multitask LearningHuaizheng Zhang, Yong Luo, Qiming Ai, Yonggang Wen 等ACM MM 2020 · 被引用 17 次
