TM2D: Bimodality Driven 3D Dance Generation via Music-Text Integration
Kehong Gong, Dongze Lian, Heng Chang, Chuan Guo, Zihang Jiang, Xinxin Zuo, Michael Bi Mi, Xinchao Wang
摘要
We propose a novel task for generating 3D dance movements that simultaneously incorporate both text and music modalities. Unlike existing works that generate dance movements using a single modality such as music, our goal is to produce richer dance movements guided by the instructive information provided by the text. However, the lack of paired motion data with both music and text modalities limits the ability to generate dance movements that integrate both. To alleviate this challenge, we propose to utilize a 3D human motion VQ-VAE to project the motions of the two datasets into a latent space consisting of quantized vectors, which effectively mix the motion tokens from the two datasets with different distributions for training. Additionally, we propose a cross-modal transformer to integrate text instructions into motion generation architecture for generating 3D dance movements without degrading the performance of music-conditioned dance generation. To better evaluate the quality of the generated motion, we introduce two novel metrics, namely Motion Prediction Distance (MPD) and Freezing Score (FS), to measure the coherence and freezing percentage of the generated motion. Extensive experiments show that our approach can generate realistic and coherent dance movements conditioned on both text and music while maintaining comparable performance with the two single modalities. Code is available at https://garfield-kh.github.io/TM2D/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Priority-Centric Human Motion Generation in Discrete Latent SpaceHanyang Kong, Kehong Gong, Dongze Lian, Michael Bi Mi 等ICCV 2023 · 被引用 81 次
- Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance AccompanimentLi Siyao, Tianpei Gu, Zhitao Yang, Zhengyu Lin 等ICLR 2024 · 被引用 54 次
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao 等SIGGRAPH 2024 · 被引用 39 次
- Generative Human Motion Stylization in Latent SpaceChuan Guo, Yuxuan Mu, Xinxin Zuo, Peng Dai 等ICLR 2024 · 被引用 30 次
- Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion GenerationBohong Chen, Yumeng Li, Yao-Xiang Ding, Tianjia Shao 等ACM MM 2024 · 被引用 26 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 被引用 701 次
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 被引用 672 次
相关 Paper
- UniMuMo: Unified Text, Music, and Motion GenerationHan Yang, Kun Su, Yutong Zhang, Jiaben Chen 等AAAI 2025 · 被引用 3 次
- MDD: A Dataset for Text-and-Music Conditioned Duet Dance GenerationPrerit Gupta, Jason Alexander Fotso-Puepi, Zhengyuan Li, Jay Mehta 等ICCV 2025 · 被引用 1 次
- MotivDance: Fine-Grained Text-Guided Motivation Choreography with Music SynchronizationChenguang Li, Yu-Hui Wen, Liping JingAAAI 2026
- Walk Before You Dance: High-fidelity and Editable Dance Synthesis via Generative Masked Motion PriorForam Niravbhai Shah, Parshwa Shah, Muhammad Usama Saleem, Ekkasit Pinyoanuntapong 等AAAI 2026 · 被引用 4 次
- AttT2M: Text-Driven Human Motion Generation with Multi-Perspective Attention MechanismChongyang Zhong, Lei Hu, Zihao Zhang, Shihong XiaICCV 2023 · 被引用 127 次
