Temporally Guided Music-to-Body-Movement Generation
Hsuan-Kai Kao, Li Su
Abstract
This paper presents a neural network model to generate virtual violinistâĂŹs 3-D skeleton movements from music audio. Improved from the conventional recurrent neural network models for generating 2-D skeleton data in previous works, the proposed model incorporates an encoder-decoder architecture, as well as the selfattention mechanism to model the complicated dynamics in body movement sequences. To facilitate the optimization of self-attention model, beat tracking is applied to determine effective sizes and boundaries of the training examples. The decoder is accompanied with a refining network and a bowing attack inference mechanism to emphasize the right-hand behavior and bowing attack timing. Both objective and subjective evaluations reveal that the proposed model outperforms the state-of-the-art methods. To the best of our knowledge, this work represents the first attempt to generate 3-D violinistsâĂŹ body movements considering key features in musical body movement.
• Computing methodologies → Motion processing; • Applied computing → Media arts; Sound and music computing; • Humancentered computing → Sound-based input / output.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbddd155-3853-405c-933a-9a9ddc390d6eCited by top-tier papers9
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic MemoryLi Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin et al.CVPR 2022 · 170 citations
- TM2D: Bimodality Driven 3D Dance Generation via Music-Text IntegrationKehong Gong, Dongze Lian, Heng Chang, Chuan Guo et al.ICCV 2023 · 103 citations
- Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion ModelYin Wang, Zhiying Leng, Frederick W. B. Li, Shun-Cheng Wu et al.ICCV 2023 · 95 citations
- Audio Matters Too! Enhancing Markerless Motion Capture with Audio Signals for String Performance CaptureYitong Jin, Zhiping Qiu, Yi Shi, Shuangpeng Sun et al.SIGGRAPH 2024 · 9 citations
Related papers
- A Human-Computer Duet System for Music PerformanceYuen-Jen Lin, Hsuan-Kai Kao, Yih-Chih Tseng, Ming Tsai et al.ACM MM 2020 · 6 citations
- SelfTalk: A Self-Supervised Commutative Training Diagram to Comprehend 3D Talking FacesZiqiao Peng, Yihao Luo, Yue Shi, Hao Xu et al.ACM MM 2023 · 56 citations
- EchoAvatar: Real-time Generative Avatar Animation from Audio StreamsBohong Chen, Yumeng Li, Yinglin Xu, Youyi Zheng et al.SIGGRAPH 2026
- Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational AutoencodersJing Li, Di Kang, Wenjie Pei, Xuefei Zhe et al.ICCV 2021 · 144 citations
- Self-supervised Dance Video Synthesis Conditioned on MusicXuanchi Ren, Haoran Li, Zijian Huang, Qifeng ChenACM MM 2020 · 68 citations
