DialoGen: Towards Dialog Gesture Generation via Identity-Decoupled Style Guidance in Interactive Diffusion Model
Weiyu Zhao, Chenyang Wang, Liangxiao Hu, Zonglin Li, Wei Yu, Shengping Zhang
Abstract
We propose DialoGen, a novel framework for generating realistic gestures for both interlocutors in dialog scenarios, conditioned on conversational audios. Unlike most existing methods that focus solely on a single speaker, DialoGen simultaneously generates synchronized gestures for both participants while also embedding identity-decoupled style into generated gestures that enhance realism and expressiveness. To ensure precise synchronization between interlocutors, DialoGen adopts an interactive dual-diffusion model with mutual interaction estimation, which integrates interaction correlation into the diffusion process. More importantly, by leveraging supervised contrastive learning, we develop the identity-decoupled style guidance to adaptively decompose the identity-specific style of interlocutors into latent space, enabling multi-style dialog gesture generation. Extensive experimental results demonstrate that our model significantly outperforms existing methods in generating realistic, speech-aligned, identity-specific gestures, offering a high-quality solution for various dialog scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Human Motion Diffusion as a Generative PriorYoni Shafir, Guy Tevet, Roy Kapon, Amit Haim BermanoICLR 2024 · 371 citations
- GestureDiffuCLIP: Gesture Diffusion Model with CLIP LatentsTenglong Ao, Zeyi Zhang, Libin LiuSIGGRAPH 2023 · 151 citations
- Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational AutoencodersJing Li, Di Kang, Wenjie Pei, Xuefei Zhe et al.ICCV 2021 · 144 citations
- Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and SynthesisGilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori et al.ICCV 2019 · 114 citations
Related papers
- DIDiffGes: Decoupled Semi-Implicit Diffusion Models for Real-time Gesture Generation from SpeechYongkang Cheng, Shaoli Huang, Xuelin Chen, Jifeng Ning et al.AAAI 2025 · 3 citations
- LivelySpeaker: Towards Semantic-Aware Co-Speech Gesture GenerationYihao Zhi, Xiaodong Cun, Xuelin Chen, Xi Shen et al.ICCV 2023 · 48 citations
- Emotional Speech-Driven 3D Body Animation via Disentangled Latent DiffusionKiran Chhatre, Radek Danecek, Nikos Athanasiou, Giorgio Becherini et al.CVPR 2024
- Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion ModelXu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin et al.CVPR 2024
- DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture GenerationYICHEN PENG, Jyun-Ting Song, Siyeol Jung, RUOFAN LIU et al.CVPR 2026 · 7 citations
