EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling
Haiyang Liu, Zihao Zhu, Giorgio Becherini, Yichen Peng, Mingyang Su, You Zhou, Xuefei Zhe, Naoya Iwamoto, Bo Zheng, Michael J. Black
Abstract
We propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements. To achieve this, we first introduce BEAT2 (BEAT-SMPLX-FLAME), a new mesh-level holistic co-speech dataset. BEAT2 combines a MoShed SMPL-X body with FLAME head parameters and further refines the modeling of head, neck, and finger movements, offering a community-standardized, high-quality 3D motion captured dataset. EMAGE leverages masked body gesture priors during training to boost inference performance. It involves a Masked Audio Gesture Transformer, facilitating joint training on audio-to-gesture generation and masked gesture reconstruction to effectively encode audio and body gesture hints. Encoded body hints from masked gestures are then separately employed to generate facial and body movements. Moreover, EMAGE adaptively merges speech features from the audio's rhythm and content and utilizes four compositional VQ-VAEs to enhance the results' fidelity and diversity. Experiments demonstrate that EMAGE generates holistic gestures with state-of-the-art performance and is flexible in accepting predefined spatial-temporal gesture inputs, generating complete, audio-synchronized results. Our code and dataset are available. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers49
- GestureLSM: Latent Shortcut Based Co-Speech Gesture Generation with Spatial-Temporal ModelingPinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu et al.ICCV 2025 · 54 citations
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao et al.SIGGRAPH 2024 · 39 citations
- Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion GenerationBohong Chen, Yumeng Li, Yao-Xiang Ding, Tianjia Shao et al.ACM MM 2024 · 26 citations
- The Quest for Generalizable Motion Generation: Data, Model, and EvaluationJing Lin, Ruisi Wang, Junzhe Lu, Ziqi Huang et al.ICLR 2026 · 23 citations
- MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal ControlsYuxuan Bian, Ailing Zeng, Xuan Ju, Xian Liu et al.AAAI 2025 · 22 citations
Builds on31
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
Related papers
- LiveGesture: Streamable Co-Speech Gesture Generation ModelMuhammad Usama Saleem, Mayur Jagdishbhai Patel, Ekkasit Pinyoanuntapong, Zhongxing Qin et al.CVPR 2026 · 4 citations
- From Audio to Photoreal Embodiment: Synthesizing Humans in ConversationsEvonne Ng, Javier Romero, Timur M. Bagautdinov, Shaojie Bai et al.CVPR 2024 · 37 citations
- Democratizing High-Fidelity Co-Speech Gesture Video GenerationXu Yang, Shaoli Huang, Shenbo Xie, Xuelin Chen et al.ICCV 2025 · 1 citation
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with TransformerKunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost et al.SIGGRAPH 2023 · 18 citations
- PyraMotion: Attentional Pyramid-Structured Motion Integration for Co-Speech 3D Gesture SynthesisZhizhuo Yin, Yuk Hang Tsui, Pan HuiNeurIPS 2025 · 6 citations
