World-Model Inspired Emotion-aware Token Refinement for Training-Free Multimodal Emotion Recognition
Kejun Liu, Yuanyuan Liu, Ke Wang, Zhe Chen, Yibing Zhan, Wei Xiang, Hongyan Zhang
Abstract
Multimodal Large Language Models (MLLMs) show promise for Multimodal Emotion Recognition (MER) but often remain unreliable because sparse emotional cues could be easily overwhelmed and affected by redundant context. While fine-tuning is effective, it is usually costly when using large models. Training-free methods like chain-of-thought reasoning provide a practical alternative, but they mostly rely on heuristic prompting to influence the model behaviors and do not explicitly focus on emotion relevant tokens internally, which would allow decision-relevant emotional tokens to be diluted by environmental noise, resulting in unstable predictions. To address this limitation without training, we rethink MER from a world-model perspective that treats emotion as a latent state inferred from noisy and redundant multimodal observations. Under frozen parameters, this view suggests that robustness depends on constraining why and how tokens contribute to inference. Based on this insight, we propose WETR (World-Model inspired Emotion-aware Token Refinement), a training-free, plug-and-play regulator that reshapes token usage through two mechanisms: Noise-suppressed Token Selection (NTS), which suppresses redundant intra-modal noise, and State-strengthened Token Reweighting (STR), which amplifies decision-relevant emotional tokens. Experiments on multiple MER benchmarks demonstrate that WETR consistently improves accuracy and stability under frozen parameters, which also improves token-level interpretability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningZebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang et al.NeurIPS 2024 · 293 citations
- MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the WildYuanyuan Liu, Wei Dai, Chuanxu Feng, Wenbin Wang et al.ACM MM 2022 · 83 citations
- Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion ReasoningZhiyuan Han, Beier Zhu, Yanlong Xu, Peipei Song et al.ACM MM 2025 · 7 citations
- Visual Prompting in LLMs for Enhancing Emotion RecognitionQixuan Zhang, Zhifeng Wang, Dylan Zhang, Wenjia Niu et al.EMNLP 2024 · 5 citations
- Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic TasksZhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su et al.CVPR 2024
Related papers
- Catch Your Emotion: Sharpening Emotion Perception in Multimodal Large Language ModelsYiyang Fang, Jian Liang, Wenke Huang, He Li et al.ICML 2025
- ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language ModelsMingrui Wu, Xinyue Cai, Jiayi Ji, Jiale Li et al.NeurIPS 2024 · 50 citations
- SeLaR: Selective Latent Reasoning in Large Language ModelsRenyu Fu, Guibo LuoACL 2026 · 2 citations
- Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context LearningYanshu Li, Jianjiang Yang, Ziteng Yang, Bozheng Li et al.AAAI 2026 · 9 citations
- EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion ReasoningDingdong WANG, Shujie LIU, Tianhua Zhang, Youjun Chen et al.ICLR 2026 · 22 citations
