Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
Zebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang, Yuxiang Lin, Zheng Lian, Xiaojiang Peng, Alexander G. Hauptmann
Abstract
Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, existing Multimodal Large Language Models (MLLMs) face challenges in integrating audio and recognizing subtle facial micro-expressions. To address this, we introduce the MERR dataset, containing 28,618 coarse-grained and 4,487 fine-grained annotated samples across diverse emotional categories. This dataset enables models to learn from varied scenarios and generalize to real-world applications. Furthermore, we propose Emotion-LLaMA, a model that seamlessly integrates audio, visual, and textual inputs through emotion-specific encoders. By aligning features into a shared space and employing a modified LLaMA model with instruction tuning, Emotion-LLaMA significantly enhances both emotional recognition and reasoning capabilities. Extensive evaluations show Emotion-LLaMA outperforms other MLLMs, achieving top scores in Clue Overlap (7.83) and Label Overlap (6.25) on EMER, an F1 score of 0.9036 on MER2023-SEMI challenge, and the highest UAR (45.59) and WAR (59.37) in zero-shot evaluations on DFEW dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44c992ed-0c04-44af-875c-54ed9019b5d9Cited by top-tier papers53
- MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language ModelsFan Zhang, Zebang Cheng, Chong Deng, Haoxuan Li et al.ICLR 2026 · 23 citations
- Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective ModelFuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng et al.ACM MM 2024 · 13 citations
- VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation ModelsZhicheng Zhang, Weicheng Wang, Yongjie Zhu, Wenyu Qin et al.NeurIPS 2025 · 11 citations
- PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector AlgebraXiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang et al.ICLR 2026 · 11 citations
- EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?Zheng Lian, Licai Sun, Lan Chen, Haoyu Chen et al.ICLR 2026 · 10 citations
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- DEEMO: De-identity Multimodal Emotion Recognition and ReasoningDeng Li, Bohao Xing, Xin Liu, Baiqiang Xia et al.ACM MM 2025 · 8 citations
- AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language ModelsZheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun et al.ICML 2025
- FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and ReasoningZhuozhao Hu, Kaishen Yuan, Xin Liu, Zitong Yu et al.ACM MM 2025 · 3 citations
- Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion ReasoningZhiyuan Han, Beier Zhu, Yanlong Xu, Peipei Song et al.ACM MM 2025 · 7 citations
- Causal-ERC: A Multimodal Framework with Causal Prompting for Emotion Recognition in Conversations with Large Language ModelsRan Jing, Geng Tu, Yice Zhang, Ruifeng XuAAAI 2026
