PromptEmo: Learning Emotion with Bilateral Textual Prompts in Multi-Domain Open-set Scenarios
Xinyi Zeng, Yuxiang Yang, Pinxian Zeng, Wenxia Yin, Bo Liu, Xi Wu, Yan Wang
Abstract
Facial Expression Recognition (FER) is crucial to human-computer interaction. Existing cross-domain FER (CD-FER) methods mainly focus on single-source closed-set scenarios, transferring knowledge from a single source domain to a target domain with identical class sets. However, CD-FER faces two real-world challenges: 1) the need to leverage information from multiple sources, leading to multi-domain shift, and 2) the necessity to recognize unseen target classes, resulting in class shift. These issues give rise to a novel and challenging task, which we define as Multi-domain Open-set FER (MO-FER). In this paper, we propose PromptEmo, a novel CLIP-based framework that leverages bilateral textual prompts to address both shifts in the MO-FER task. Leveraging the generalizability of LLM, PromptEmo constructs trainable positive prompts with LLM-generated emotion descriptions for seen classes, as well as template-derived negative prompts to enhance the reasoning for unseen classes. Then, we introduce a modal-task optimization paradigm organized from two perspectives: textual semantics and visual domains, yielding Intra-modal Space-specific Optimization (ISO) and Cross-modal Emotion-aware Interaction (CEI) strategies. ISO refines the CLIP-based textual space to ensure semantic separation between bilateral prompts and improves the latent visual space by promoting inter-domain alignment. Founded on ISO, CEI facilitates effective vision-language interactions, resulting in four joint loss terms that improve emotion recognition by shaping a domain-invariant, discriminative feature space. PromptEmo surpasses the current SOTA method by 7.7% AUC on unseen classes across four FER datasets, serving as a strong baseline for the MO-FER task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Former-DFER: Dynamic Facial Expression Recognition TransformerZengqun Zhao, Qingshan LiuACM MM 2021 · 185 citations
- Progressive Graph Learning for Open-Set Domain AdaptationYadan Luo, Zijian Wang, Zi Huang, Mahsa BaktashmotlaghICML 2020 · 114 citations
- CLIPCEIL: Domain Generalization through CLIP via Channel rEfinement and Image-text aLignmentXi Yu, Shinjae Yoo, Yuewei LinNeurIPS 2024 · 36 citations
Related papers
- Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive PromptingYuanyuan Liu, Yuxuan Huang, Shuyang Liu, Yibing Zhan et al.ACM MM 2024 · 15 citations
- Multimodal Prompt Alignment for Facial Expression RecognitionFuyan Ma, Yiran He, Bin Sun, Shutao LiICCV 2025 · 5 citations
- Learning with Alignments: Tackling the Inter- and Intra-domain Shifts for Cross-multidomain Facial Expression RecognitionYuxiang Yang, Lu Wen, Xinyi Zeng, Yuanyuan Xu et al.ACM MM 2024 · 7 citations
- Domain Knowledge Enhanced Vision-Language Pretrained Model for Dynamic Facial Expression RecognitionLiupeng Li, Yuhua Zheng, Shupeng Liu, Xiaoyin Xu et al.ACM MM 2024 · 4 citations
- Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum LearningXinran Li, Yu Liu, Jiaqi Qiao, Xiujuan XuAAAI 2026 · 1 citation
