Role-aware Interaction Generation from Textual Description
Mikihiro Tanaka, Kent Fujiwara
Abstract
This research tackles the problem of generating interaction between two human actors corresponding to textual description. We claim that certain interactions, which we call asymmetric interactions, involve a relationship between an actor and a receiver, whose motions significantly differ depending on the assigned role. However, existing studies of interaction generation attempt to learn the correspondence between a single label and the motions of both actors combined, overlooking differences in individual roles. We consider a novel problem of role-aware interaction generation, where roles can be designated before generation. We translate the text of the asymmetric interactions into active and passive voice to ensure the textual context is consistent with each role. We propose a model that learns to generate motions of the designated role, which together form a mutually consistent interaction. As the model treats individual motions separately, it can be pretrained to derive knowledge from single-person motion data for more accurate interactions. Moreover, we introduce a method inspired by Permutation Invariant Training (PIT) that can automatically learn which of the two actions corresponds to an actor or a receiver without additional annotation. We further present cases where existing evaluation metrics fail to accurately assess the quality of generated interactions, and propose a novel metric, Mutual Consistency, to address such shortcomings. Experimental results demonstrate the efficacy of our method, as well as the necessity of the proposed metric. Our code is available at https://github.com/ line/Human-Interaction-Generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object DynamicsJiawei Gao, Ziqin Wang, Zeqi Xiao, Jingbo Wang et al.NeurIPS 2024 · 57 citations
- ReGenNet: Towards Human Action-Reaction SynthesisLiang Xu, Yizhou Zhou, Yichao Yan, Xin Jin et al.CVPR 2024 · 18 citations
- Inter-X: Towards Versatile Human-Human Interaction AnalysisLiang Xu, Xintao Lv, Yichao Yan, Xin Jin et al.CVPR 2024 · 18 citations
- Causal Motion Diffusion Models for Autoregressive Motion GenerationQing Yu, Akihisa Watanabe, Kent FujiwaraCVPR 2026 · 9 citations
- ProjFlow: Projection Sampling with Flow Matching for Zero‑Shot Exact Spatial Motion ControlAkihisa Watanabe, Qing Yu, Edgar Simo-Serra, Kent FujiwaraCVPR 2026 · 6 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 1,527 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
Related papers
- Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction GenerationQingxuan Wu, Zhiyang Dou, chuan guo, Yiming Huang et al.ICLR 2026 · 10 citations
- InterSyn: Interleaved Learning for Dynamic Motion Synthesis in the WildYiyi Ma, Yuanzhi Liang, Xiu Li, Chi Zhang et al.ICCV 2025 · 4 citations
- Dual Reciprocal Learning of Language-based Human Motion Understanding and GenerationChen Liang, Zhicheng Shi, Wenguan Wang, Yi YangICCV 2025 · 1 citation
- TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion GenerationYabiao Wang, Shuo Wang, Jiangning Zhang, Ke Fan et al.CVPR 2025
- Customizing Text-to-Image Generation with Inverted InteractionMengmeng Ge, Xu Jia, Takashi Isobe, Xiaomin Li et al.ACM MM 2024 · 3 citations
