ELGAR: Expressive Cello Performance Motion Generation for Audio Rendition
Zhiping Qiu, Yitong Jin, Yuan Wang, Yi Shi, Chao Tan, Chongwu Wang, Xiaobing Li, Feng Yu, Tao Yu, Qionghai Dai
Abstract
The art of instrument performance stands as a vivid manifestation of human creativity and emotion. Nonetheless, generating instrument performance motions is a highly challenging task, as it requires not only capturing intricate movements but also reconstructing the complex dynamics of the performer-instrument interaction. While existing works primarily focus on modeling partial body motions, we propose Expressive ceLlo performance motion Generation for Audio Rendition (ELGAR), a state-of-the-art diffusion-based framework for whole-body fine-grained instrument performance motion generation solely from audio. To emphasize the interactive nature of the instrument performance, we introduce Hand Interactive Contact Loss (HICL) and Bow Interactive Contact Loss (BICL), which effectively guarantee the authenticity of the interplay. Moreover, to better evaluate whether the generated motions align with the semantic context of the music audio, we design novel metrics specifically for string instrument performance motion generation, including finger-contact distance, bow-string distance, and bowing score. Extensive evaluations and ablation studies are conducted to validate the efficacy of the proposed methods. In addition, we put forward a motion generation dataset SPD-GEN, collated and normalized from the MoCap dataset SPD. As demonstrated, ELGAR has shown great potential in generating instrument performance motions with complicated and fast interactions, which will promote further development in areas such as animation, music education, interactive art creation, etc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- MUSIC: Learning Muscle-Driven Dexterous Hand ControlPei Xu, Yufei Ye, Shuchun Sun, Yu Ding et al.SIGGRAPH 2026
- EchoAvatar: Real-time Generative Avatar Animation from Audio StreamsBohong Chen, Yumeng Li, Yinglin Xu, Youyi Zheng et al.SIGGRAPH 2026
Builds on24
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- Audio Matters Too! Enhancing Markerless Motion Capture with Audio Signals for String Performance CaptureYitong Jin, Zhiping Qiu, Yi Shi, Shuangpeng Sun et al.SIGGRAPH 2024 · 9 citations
- HandDiffuse: Generative Controllers for Two-Hand Interactions via Diffusion ModelsPei LinAAAI 2025 · 1 citation
- 🎧MOSPA: Human Motion Generation Driven by Spatial AudioShuyang Xu, Zhiyang Dou, Mingyi Shi, Liang Pan et al.NeurIPS 2025 · 13 citations
- SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction SynthesisWenkun He, Yun Liu, Ruitao Liu, Li YiICCV 2025 · 1 citation
- AvatarGO: Zero-shot 4D Human-Object Interaction Generation and AnimationYukang Cao, Liang Pan, Kai Han, Kwan-Yee K. Wong et al.ICLR 2025
