EmIT: Emotional Interaction control in Text-to-image diffusion models
Haofan Zhang, Shangfei Wang
摘要
Although current work of text-to-Image generation can preliminarily generate images from the descriptions of human-object interactions, it fails to consider the emotions involved in human-object interactions. While people often experience emotions when using objects or interacting with them. Therefore, in this paper, we propose Emotional Interaction Generation task, a novel image generation task, which generates emotionally expressive human-object interaction images from given prompts, human-object interaction (HOI) region, and emotions. First, we construct a new emotional interaction dataset, called EmotionHOI, which including 47,776 images with content prompt, emotions and human-object interaction bounding box. Second, we propose an emotion-aware text-to-image diffusion model, named EmIT, for emotional interaction generation. Specifically, EmIT consists of three components: (1) an emotion interaction tokenizer that encodes subject, object, action, and emotion into structured tokens; (2) an Emo-Interaction Self-Attention that preliminarily guides the latent space to conduct hybrid learning with emotional interaction tokens; and (3) a Hierarchical Emotion-Visual Cross-Attention that further focus on grounding affect-such as pose, gaze, or interaction intensity-into specific spatial regions and capture subtle emotional variations. These components jointly model interaction semantics and emotional context, enabling EmIT to generate images that are both behaviorally coherent and emotionally expressive. Experimental results on the EmotionHOI dataset demonstrate the superiority of the proposed model.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion ModelsLiulei Li, Wenguan Wang, Yi YangNeurIPS 2024 · 被引用 29 次
- Semantic-Aware Human Object Interaction Image GenerationZhu Xu, Qingchao Chen, Yuxin Peng, Yang LiuICML 2024 · 被引用 9 次
- Learning to Generate Human-Human-Object Interactions from Textual DescriptionsJeonghyeon Na, Sangwon Baik, Inhee Lee, Junyoung Lee 等NeurIPS 2025 · 被引用 3 次
- Make Me Happier: Evoking Emotions through Image Diffusion ModelsQing Lin, Jingfeng Zhang, Yew-Soon Ong, Mengmi ZhangICCV 2025 · 被引用 4 次
- Visual Relation Diffusion for Human-Object Interaction DetectionPing Cao, Yepeng Tang, Chunjie Zhang, Xiaolong Zheng 等ICCV 2025 · 被引用 1 次
