Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
Xu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin, Zhiyong Wu, Sicheng Yang, Minglei Li, Zhiyi Chen, Songcen Xu, Xiaofei Wu
摘要
Co-speech gestures, if presented in the lively form of videos, can achieve superior visual effects in humanmachine interaction. While previous works mostly generate structural human skeletons, resulting in the omission of appearance information, we focus on the direct generation of audio-driven co-speech gesture videos in this work. There are two main challenges: 1) A suitable motion feature is needed to describe complex human movements with crucial appearance information. 2) Gestures and speech exhibit inherent dependencies and should be temporally aligned even of arbitrary length. To solve these problems, we present a novel motion-decoupled framework to generate co-speech gesture videos. Specifically, we first introduce a well-designed nonlinear TPS transformation to obtain latent motion features preserving essential appearance information. Then a transformer-based diffusion model is proposed to learn the temporal correlation between gestures and speech, and performs generation in the latent motion space, followed by an optimal motion selection module to produce long-term coherent and consistent gesture videos. For better visual perception, we further design a refinement network focusing on missing details of certain areas. Extensive experimental results show that our proposed framework significantly outperforms existing ap-proaches in both motion and video-related evaluations. Our code, demos, and more resources are available at https: //github.com/thuhcsi/S2G-MDDiffusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- MagicFight: Personalized Martial Arts Combat Video GenerationJiancheng Huang, Mingfu Yan, Songyan Chen, Yi Huang 等ACM MM 2024 · 被引用 16 次
- MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative RefinementXu He, Zhiyong Wu, Xiaoyu Li, Di Kang 等AAAI 2025 · 被引用 11 次
- EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion GenerationXiangyue Zhang, Jianfang Li, Jiaxu Zhang, Jianqiang Ren 等ACM MM 2025 · 被引用 7 次
- Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level ConstraintsXiangyue Zhang, Jianfang Li, Jianqiang Ren, Jiaxu ZhangAAAI 2026 · 被引用 7 次
- Video Motion GraphsHaiyang Liu, Zhan Xu, Fa-Ting Hong, Hsin-Ping Huang 等ICCV 2025 · 被引用 6 次
它引用的顶会 Paper25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 被引用 1,527 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 被引用 701 次
相关 Paper
- Taming Diffusion Models for Audio-Driven Co-Speech Gesture GenerationLingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian 等CVPR 2023
- MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture GenerationXiaofeng Mao, Zhengkai Jiang, Qilin Wang, Chencan Fu 等ACM MM 2024 · 被引用 7 次
- Co-Speech Gesture Video Generation with Implicit Motion-Audio EntanglementXinjie Li, Ziyi Chen, Xinlu Yu, Iek-Heng Chu 等CVPR 2025
- Democratizing High-Fidelity Co-Speech Gesture Video GenerationXu Yang, Shaoli Huang, Shenbo Xie, Xuelin Chen 等ICCV 2025 · 被引用 1 次
- Audio-Driven Co-Speech Gesture Video GenerationXian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du 等NeurIPS 2022 · 被引用 77 次
