GestureHYDRA: Semantic Co-Speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
Quanwei Yang, Luying Huang, Kaisiyuan Wang, Jiazhi Guan, Shengyi He, Fengguo Li, Hang Zhou, Lingyun Yu, Yingying Li, Haocheng Feng, Hongtao Xie
摘要
While increasing attention has been paid to co-speech gesture synthesis, most previous works neglect to investigate hand gestures with explicit and essential semantics. In this paper, we study co-speech gesture generation with an emphasis on specific hand gesture activation, which can deliver more instructional information than common body movements. To achieve this, we first build a high-quality dataset of 3D human body movements including a set of semantically explicit hand gestures that are commonly used by live streamers. Then we present a hybrid-modality gesture generation system GestureHYDRA built upon a hybridmodality diffusion transformer architecture with novelly designed motion-style injective transformer layers, which enables advanced gesture modeling ability and versatile gesture operations. To guarantee these specific hand gestures can be activated, we introduce a cascaded retrieval-augmented generation strategy built upon a semantic gesture repository annotated for each subject and an adaptive audio-gesture synchronization mechanism, which substantially improves semantic gesture activation and production efficiency. Quantitative and qualitative experiments demonstrate that our proposed approach achieves superior performance over all the counterparts. The project page can be found https://mumuwei.github.io/GestureHYDRA/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Multi-level Causal LLM-based Text-to-Motion Generation with Human AlignmentChen Xiaodong, Qian Bao, Xudong Liu, Jianping Fang 等CVPR 2026
- EchoAvatar: Real-time Generative Avatar Animation from Audio StreamsBohong Chen, Yumeng Li, Yinglin Xu, Youyi Zheng 等SIGGRAPH 2026
它引用的顶会 Paper33
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- DiffusionDet: Diffusion Model for Object DetectionShoufa Chen, Peize Sun, Yibing Song, Ping LuoICCV 2023 · 被引用 715 次
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 被引用 701 次
相关 Paper
- MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture GenerationXiaofeng Mao, Zhengkai Jiang, Qilin Wang, Chencan Fu 等ACM MM 2024 · 被引用 7 次
- DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-Driven Holistic 3D Expression and Gesture GenerationJunming Chen, Yunfei Liu, Jianan Wang, Ailing Zeng 等CVPR 2024
- SemGesture: Synthesizing Semantically Enhanced and Coherent GesturesPengsheng Liu, Zhaojie Chu, Xiaofen Xing, Xiangmin XuACM MM 2025 · 被引用 2 次
- Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion ModelXu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin 等CVPR 2024
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with TransformerKunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost 等SIGGRAPH 2023 · 被引用 18 次
