Learning Emotion Representations from Verbal and Nonverbal Communication
Sitao Zhang, Yimu Pan, James Z. Wang
摘要
Emotion understanding is an essential but highly challenging component of artificial general intelligence. The absence of extensive annotated datasets has significantly impeded advancements in this field. We present Emotion-CLIP, the first pre-training paradigm to extract visual emotion representations from verbal and nonverbal communication using only uncurated data. Compared to numerical labels or descriptions used in previous methods, communication naturally contains emotion information. Furthermore, acquiring emotion representations from communication is more congruent with the human learning process. We guide EmotionCLIP to attend to nonverbal emotion cues through subject-aware context encoding and verbal emotion cues using sentiment-guided contrastive learning. Extensive experiments validates the effectiveness and transferability of EmotionCLIP. Using merely linear-probe evaluation protocol, EmotionCLIP outperforms the state-of-theart supervised visual emotion recognition methods and rivals many multimodal approaches across various benchmarks. We anticipate that the advent of EmotionCLIP will address the prevailing issue of data scarcity in emotion understanding, thereby fostering progress in related domains. The code and pre-trained models are available at https://github.com/Xeaver/EmotionCLIP . * equal contribution Conversation: -Dad. Who is Dede? -Jesus. She was the love of my life and I was too stupid to realized it. I lost her because of something so dumb. 0 1 1 0 0 Categorical Label: Shocked Regretful Description: An old man is chatting with his son while eating at a restaurant.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningZebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang 等NeurIPS 2024 · 被引用 293 次
- VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation ModelsZhicheng Zhang, Weicheng Wang, Yongjie Zhu, Wenyu Qin 等NeurIPS 2025 · 被引用 11 次
- Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text PairsDaiqing Wu, Dongbao Yang, Yu Zhou, Can MaACM MM 2024 · 被引用 6 次
- Visual Prompting in LLMs for Enhancing Emotion RecognitionQixuan Zhang, Zhifeng Wang, Dylan Zhang, Wenjia Niu 等EMNLP 2024 · 被引用 5 次
- Two in One Go: Single-stage Emotion Recognition with Decoupled Subject-context TransformerXinpeng Li, Teng Wang, Jian Zhao, Shuyi Mao 等ACM MM 2024 · 被引用 3 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- Emotion in a Bottle: Information Bottleneck Guided Disentanglement for Emotion Domain AdaptationJiankun Zhu, Sicheng Zhao, Lulu Tian, Jing Jiang 等ACM MM 2025
- LES-CLIP: A Lightweight Emotion-Sensitive Adaptation of CLIP for Precise Similar Emotion DiscriminationXiao Fu, Pengyu Wang, Wei Xi, Kun Zhao 等ACM MM 2025 · 被引用 2 次
- VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionHao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu 等ACM MM 2025 · 被引用 9 次
- Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive PromptingYuanyuan Liu, Yuxuan Huang, Shuyang Liu, Yibing Zhan 等ACM MM 2024 · 被引用 15 次
- Grounding Emotion Recognition with Visual Prototypes: VEGA - Revisiting CLIP in MERCGuanyu Hu, Dimitrios Kollias, Xinyu YangACM MM 2025 · 被引用 5 次
