EmoSym: A Symbiotic Framework for Unified Emotional Understanding and Generation via Latent Reasoning
Yijie Zhu, Yibo Lyu, Zitong Yu, Rui Shao, Kaiyang Zhou, Liqiang Nie
Abstract
Current affective computing paradigms often treat emotional understanding and generation as separate tasks, yet they inherently possess symbiotic potential for mutual enhancement. In this paper, we aim to bridge the gap by developing a unified framework. The primary challenge lies in the extraction of precise and semantically rich representations of abstract emotions, which are crucial for both tasks. To address this, we harness the Chain-of-Thought reasoning at the latent space of multimodal large language models and propose EmoSym, a unified framework built upon this advanced foundation. Our framework is executed through three key steps: 1) Emotional reasoning knowledge compression. To enable efficient transfer of emotional reasoning priors, we design specialized reasoning tokens to compact emotion-aware contexts from external reasoning knowledge bases into latent representations. 2) Verifiable reinforcement reasoning optimization. To ensure more reliable and consistent emotional reasoning, we develop a verifiable reinforcement learning paradigm to further enhance the reasoning token by emotion-specific verifiable reward signals. Processed through the above two steps, the reasoning token simultaneously enhances emotional understanding while enriching semantic representations, benefiting subsequent emotional generation tasks. 3) Reasoning-augmented generation and online feedback. We then fuse it with emotional representations and feed them into a diffusion model to generate emotion-evoking images. Additionally, to create a generative-to-understanding enhancement feedback, we propose an Online Emotional Memory Bank (OEMB). It leverages newly generated images to progressively update the training dataset in the training process to reinforce understanding. Extensive experiments demonstrate the superior capabilities of our framework in both emotional understanding and generation tasks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 25dd4bb6-aa00-4a96-bf0c-362c98089cc4Cited by top-tier papers6
- SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic ManipulationWei Li, Renshan Zhang, Rui Shao, Zhijian Fang et al.AAAI 2026 · 13 citations
- ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic ManipulationWei Li, Jizhihui Liu, Yixing Li, Junwen Tong et al.CVPR 2026 · 8 citations
- H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic ManipulationYijie Zhu, Rui Shao, Ziyang Liu, Jie He et al.AAAI 2026 · 5 citations
- GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian SplattingLei Yao, Yi Wang, Yi Zhang, Moyun Liu et al.ACM MM 2025 · 2 citations
- Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent PurificationJiayu Zhang, Shuo Ye, Qilang Ye, Zihan Song et al.ACL 2026 · 2 citations
Related papers
- Consensus-Driven Multi-Agent Cognitive Reasoning for Enhancing the Emotional Intelligence of Large Language ModelsGeng Tu, Dingming Li, Jun Huang, Ruifeng XuAAAI 2026
- EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language ModelsYiyang Fang, Wenke Huang, Pei Fu, Yihao Yang et al.CVPR 2026 · 4 citations
- HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor GenerationJiajun Zhang, Shijia Luo, Ruikang Zhang, Qi SuCVPR 2026 · 4 citations
- CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content GenerationKaishen Yuan, Yuting Zhang, Shang Gao, Yijie Zhu et al.ICLR 2026 · 10 citations
- UniMo: Unified Motion Generation and Understanding with Chain of ThoughtGuocun Wang, Kenkun Liu, Jing Lin, Guorui Song et al.AAAI 2026
