Role-aware Interaction Generation from Textual Description
Mikihiro Tanaka, Kent Fujiwara
摘要
This research tackles the problem of generating interaction between two human actors corresponding to textual description. We claim that certain interactions, which we call asymmetric interactions, involve a relationship between an actor and a receiver, whose motions significantly differ depending on the assigned role. However, existing studies of interaction generation attempt to learn the correspondence between a single label and the motions of both actors combined, overlooking differences in individual roles. We consider a novel problem of role-aware interaction generation, where roles can be designated before generation. We translate the text of the asymmetric interactions into active and passive voice to ensure the textual context is consistent with each role. We propose a model that learns to generate motions of the designated role, which together form a mutually consistent interaction. As the model treats individual motions separately, it can be pretrained to derive knowledge from single-person motion data for more accurate interactions. Moreover, we introduce a method inspired by Permutation Invariant Training (PIT) that can automatically learn which of the two actions corresponds to an actor or a receiver without additional annotation. We further present cases where existing evaluation metrics fail to accurately assess the quality of generated interactions, and propose a novel metric, Mutual Consistency, to address such shortcomings. Experimental results demonstrate the efficacy of our method, as well as the necessity of the proposed metric. Our code is available at https://github.com/ line/Human-Interaction-Generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object DynamicsJiawei Gao, Ziqin Wang, Zeqi Xiao, Jingbo Wang 等NeurIPS 2024 · 被引用 57 次
- ReGenNet: Towards Human Action-Reaction SynthesisLiang Xu, Yizhou Zhou, Yichao Yan, Xin Jin 等CVPR 2024 · 被引用 18 次
- Inter-X: Towards Versatile Human-Human Interaction AnalysisLiang Xu, Xintao Lv, Yichao Yan, Xin Jin 等CVPR 2024 · 被引用 18 次
- Causal Motion Diffusion Models for Autoregressive Motion GenerationQing Yu, Akihisa Watanabe, Kent FujiwaraCVPR 2026 · 被引用 9 次
- ProjFlow: Projection Sampling with Flow Matching for Zero‑Shot Exact Spatial Motion ControlAkihisa Watanabe, Qing Yu, Edgar Simo-Serra, Kent FujiwaraCVPR 2026 · 被引用 6 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 被引用 1,527 次
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 被引用 701 次
相关 Paper
- Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction GenerationQingxuan Wu, Zhiyang Dou, chuan guo, Yiming Huang 等ICLR 2026 · 被引用 10 次
- InterSyn: Interleaved Learning for Dynamic Motion Synthesis in the WildYiyi Ma, Yuanzhi Liang, Xiu Li, Chi Zhang 等ICCV 2025 · 被引用 4 次
- Dual Reciprocal Learning of Language-based Human Motion Understanding and GenerationChen Liang, Zhicheng Shi, Wenguan Wang, Yi YangICCV 2025 · 被引用 1 次
- TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion GenerationYabiao Wang, Shuo Wang, Jiangning Zhang, Ke Fan 等CVPR 2025
- Customizing Text-to-Image Generation with Inverted InteractionMengmeng Ge, Xu Jia, Takashi Isobe, Xiaomin Li 等ACM MM 2024 · 被引用 3 次
