Visual Prompting in LLMs for Enhancing Emotion Recognition
Qixuan Zhang, Zhifeng Wang, Dylan Zhang, Wenjia Niu, Sabrina B. Caldwell, Tom Gedeon, Yang Liu, Zhenyue Qin
Abstract
Vision Large Language Models (VLLMs) are transforming the intersection of computer vision and natural language processing. Nonetheless, the potential of using visual prompts for emotion recognition in these models remains largely unexplored and untapped. Traditional methods in VLLMs struggle with spatial localization and often discard valuable global context. To address this problem, we propose a Set-of-Vision prompting (SoV) approach that enhances zero-shot emotion recognition by using spatial information, such as bounding boxes and facial landmarks, to mark targets precisely. SoV improves accuracy in face count and emotion categorization while preserving the enriched image context. Through a battery of experimentation and analysis of recent commercial or open-source VLLMs, we evaluate the SoV model's ability to comprehend facial expressions in natural environments. Our findings demonstrate the effectiveness of integrating spatial visual prompts into VLLMs for improving emotion recognition performance. * Equal contribution † Corresponding authors Identify and locate faces in the image. (1) Box 2 Enhance grounding capabilities by box and number (2) Box + Number 1 2 3 (3) Box + Number + Facial Landmarks Analyze facial expression by spatial relationships 3
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation ModelsZhicheng Zhang, Weicheng Wang, Yongjie Zhu, Wenyu Qin et al.NeurIPS 2025 · 11 citations
- ViKey: Enhancing Temporal Understanding in Videos via Visual PromptingYeonkyung Lee, Dayun Ju, Youngmin Kim, Seil Kang et al.CVPR 2026 · 3 citations
- Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal SettingsMd Messal Monem Miah, Adrita Anika, Xi Shi, Ruihong HuangACL 2025
- MASP: Multi-Aspect Guided Emotion Reasoning with Soft Prompt Tuning In Vision-Language ModelsSangEun Lee, Yubeen Lee, Eunil Park, Wonseok ChaeAAAI 2026
- World-Model Inspired Emotion-aware Token Refinement for Training-Free Multimodal Emotion RecognitionKejun Liu, Yuanyuan Liu, Ke Wang, Zhe Chen et al.ICML 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 262 citations
Related papers
- VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language ModelsMingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li et al.AAAI 2026
- VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic AugmentationYancheng Wang, Osama Hanna, Ruiming Xie, Xianfeng Rui et al.ICLR 2026 · 4 citations
- Visual in-Context PromptingFeng Li, Qing Jiang, Hao Zhang, Tianhe Ren et al.CVPR 2024
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger et al.NeurIPS 2023 · 63 citations
- Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual PromptingGiacomo Frisoni, Lorenzo Molfetta, Mattia Buzzoni, Gianluca MoroAAAI 2026
