EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
Hongxia Xie, Chu-Jun Peng, Yu-Wen Tseng, Hung-Jen Chen, Chan-Feng Hsu, Hong-Han Shuai, Wen-Huang Cheng
摘要
Visual Instruction Tuning represents a novel learning paradigm involving the fine-tuning of pre-trained language models using task-specific instructions. This paradigm shows promising zero-shot results in various natural language processing tasks but is still unexplored in vision emotion understanding. In this work, we focus on enhancing the model's proficiency in understanding and adhering to instructions related to emotional contexts. Initially, we identify key visual clues critical to visual emotion recognition. Subsequently, we introduce a novel GPT-assisted pipeline for generating emotion visual instruction data, effectively addressing the scarcity of annotated instruction data in this domain. Expanding on the groundwork established by InstructBLIP, our proposed EmoVIT architecture incorporates emotion-specific instruction data, leveraging the powerful capabilities of Large Language Models to enhance performance. Through extensive experiments, our model showcases its proficiency in emotion classification, adeptness in affective reasoning, and competence in comprehending humor. The comparative analysis provides a robust benchmark for Emotion Visual Instruction Tuning in the era of LLMs, providing valuable insights and opening avenues for future exploration in this domain. Our code is available at https://github.com/aimmemotion/EmoVIT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningZebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang 等NeurIPS 2024 · 被引用 293 次
- VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation ModelsZhicheng Zhang, Weicheng Wang, Yongjie Zhu, Wenyu Qin 等NeurIPS 2025 · 被引用 11 次
- Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion ReasoningZhiyuan Han, Beier Zhu, Yanlong Xu, Peipei Song 等ACM MM 2025 · 被引用 7 次
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable ApproachDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma 等ICLR 2026 · 被引用 4 次
- AVERE: Improving Audiovisual Emotion Reasoning with Preference OptimizationAshutosh Chaubey, Jiacheng Pang, Maksim Siniukov, Mohammad SoleymaniICLR 2026 · 被引用 4 次
它引用的顶会 Paper12
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- MetaMorph: Multimodal Understanding and Generation via Instruction TuningShengbang Tong, David Fan, Jiachen Zhu, Yunyang Xiong 等ICCV 2025 · 被引用 14 次
- MASP: Multi-Aspect Guided Emotion Reasoning with Soft Prompt Tuning In Vision-Language ModelsSangEun Lee, Yubeen Lee, Eunil Park, Wonseok ChaeAAAI 2026
- Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction TuningBardia Safaei, Faizan Siddiqui, Jiacong Xu, Vishal M. Patel 等CVPR 2025
- VIGC: Visual Instruction Generation and CorrectionBin Wang, Fan Wu, Xiao Han, Jiahui Peng 等AAAI 2024 · 被引用 95 次
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction TuningFuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang 等ICLR 2024 · 被引用 476 次
