Lune

ACM MM2025顶会

EmIT: Emotional Interaction control in Text-to-image diffusion models

Haofan Zhang, Shangfei Wang

2025年份

摘要

Although current work of text-to-Image generation can preliminarily generate images from the descriptions of human-object interactions, it fails to consider the emotions involved in human-object interactions. While people often experience emotions when using objects or interacting with them. Therefore, in this paper, we propose Emotional Interaction Generation task, a novel image generation task, which generates emotionally expressive human-object interaction images from given prompts, human-object interaction (HOI) region, and emotions. First, we construct a new emotional interaction dataset, called EmotionHOI, which including 47,776 images with content prompt, emotions and human-object interaction bounding box. Second, we propose an emotion-aware text-to-image diffusion model, named EmIT, for emotional interaction generation. Specifically, EmIT consists of three components: (1) an emotion interaction tokenizer that encodes subject, object, action, and emotion into structured tokens; (2) an Emo-Interaction Self-Attention that preliminarily guides the latent space to conduct hybrid learning with emotional interaction tokens; and (3) a Hierarchical Emotion-Visual Cross-Attention that further focus on grounding affect-such as pose, gaze, or interaction intensity-into specific spatial regions and capture subtle emotional variations. These components jointly model interaction semantics and emotional context, enabling EmIT to generate images that are both behaviorally coherent and emotionally expressive. Experimental results on the EmotionHOI dataset demonstrate the superiority of the proposed model.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖