TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation
Haiyang Liu, Xingchao Yang, Tomoya Akiyama, Yuantian Huang, Qiaoge Li, Shigeru Kuriyama, Takafumi Taketomi
2025Year
12Top-tier citations
Abstract
TANGO is a framework designed to generate co-speech body-gesture videos using a motion graph-based retrieval approach. It first retrieves most of the reference video clips that match the target speech audio by utilizing an implicit hierarchical audio-motion embedding space. Then, it adopts a diffusion-based interpolation network to generate the remaining transition frames and smooth the discontinuities at clip boundaries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6cf1abc-814b-46fb-8c1b-4e164e183322Cited by top-tier papers12
- GestureLSM: Latent Shortcut Based Co-Speech Gesture Generation with Spatial-Temporal ModelingPinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu et al.ICCV 2025 · 54 citations
- FloodDiffusion: Tailored Diffusion Forcing for Streaming Motion GenerationYIYI CAI, Yuhan Wu, Kunhang Li, YOU ZHOU et al.CVPR 2026 · 14 citations
- EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion GenerationXiangyue Zhang, Jianfang Li, Jiaxu Zhang, Jianqiang Ren et al.ACM MM 2025 · 7 citations
- Video Motion GraphsHaiyang Liu, Zhan Xu, Fa-Ting Hong, Hsin-Ping Huang et al.ICCV 2025 · 6 citations
- LiveGesture: Streamable Co-Speech Gesture Generation ModelMuhammad Usama Saleem, Mayur Jagdishbhai Patel, Ekkasit Pinyoanuntapong, Zhongxing Qin et al.CVPR 2026 · 4 citations
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion ModelXu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin et al.CVPR 2024
- Audio-driven Neural Gesture Reenactment with Video Motion GraphsYang Zhou, Jimei Yang, Dingzeyu Li, Jun Saito et al.CVPR 2022 · 22 citations
- Audio-Driven Co-Speech Gesture Video GenerationXian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du et al.NeurIPS 2022 · 77 citations
- Co-Speech Gesture Video Generation with Implicit Motion-Audio EntanglementXinjie Li, Ziyi Chen, Xinlu Yu, Iek-Heng Chu et al.CVPR 2025
- Taming Diffusion Models for Audio-Driven Co-Speech Gesture GenerationLingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian et al.CVPR 2023
