Audio Matters Too! Enhancing Markerless Motion Capture with Audio Signals for String Performance Capture
Yitong Jin, Zhiping Qiu, Yi Shi, Shuangpeng Sun, Chongwu Wang, Donghao Pan, Jiachen Zhao, Zhenghao Liang, Yuan Wang, Xiaobing Li, Feng Yu, Tao Yu, Qionghai Dai
摘要
In this paper, we touch on the problem of markerless multi-modal human motion capture especially for string performance capture which involves inherently subtle hand-string contacts and intricate movements. To fulfill this goal, we first collect a dataset, named String Performance Dataset (SPD), featuring cello and violin performances. The dataset includes videos captured from up to 23 different views, audio signals, and detailed 3D motion annotations of the body, hands, instrument, and bow. Moreover, to acquire the detailed motion annotations, we propose an audio-guided multi-modal motion capture framework that explicitly incorporates hand-string contacts detected from the audio signals for solving detailed hand poses. This framework serves as a baseline for string performance capture in a completely markerless manner without imposing any external devices on performers, eliminating the potential of introducing distortion in such delicate movements. We argue that the movements of performers, particularly the sound-producing gestures, contain subtle information often elusive to visual methods but can be inferred and retrieved from audio cues. Consequently, we refine the vision-based motion capture results through our innovative audio-guided approach, simultaneously clarifying the contact relationship between the performer and the instrument, as deduced from the audio. We validate the proposed framework and conduct ablation studies to demonstrate its efficacy. Our results outperform current state-of-the-art vision-based algorithms, underscoring the feasibility of augmenting visual motion capture with audio modality. To the best of our knowledge, SPD is the first dataset for musical instrument performance, covering fine-grained hand motion details in a multi-modal, large-scale collection. It holds significant implications and guidance for string instrument pedagogy, animation, and virtual concerts, as well as for both musical performance analysis and generation. Our code and SPD dataset are available at https://github.com/Yitongishere/string_performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay 等ICCV 2023 · 被引用 297 次
- Temporally Guided Music-to-Body-Movement GenerationHsuan-Kai Kao, Li SuACM MM 2020 · 被引用 40 次
- Capturing detailed deformations of moving human bodiesHe Chen, Hyojoon Park, Kutay Macit, Ladislav KavanSIGGRAPH 2021 · 被引用 30 次
- Recovering 3D Hand Mesh Sequence from a Single Blurry Image: A New Dataset and Temporal UnfoldingYeonguk Oh, JoonKyu Park, Jaeha Kim, Gyeongsik Moon 等CVPR 2023
相关 Paper
- ELGAR: Expressive Cello Performance Motion Generation for Audio RenditionZhiping Qiu, Yitong Jin, Yuan Wang, Yi Shi 等SIGGRAPH 2025 · 被引用 4 次
- 🎧MOSPA: Human Motion Generation Driven by Spatial AudioShuyang Xu, Zhiyang Dou, Mingyi Shi, Liang Pan 等NeurIPS 2025 · 被引用 13 次
- HOnnotate: A Method for 3D Annotation of Hand and Object PosesShreyas Hampali, Mahdi Rad, Markus Oberweger, Vincent LepetitCVPR 2020
- PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano PerformanceQijun Gan, Song Wang, Shengtao Wu, Jianke ZhuICLR 2025 · 被引用 1 次
- MDD: A Dataset for Text-and-Music Conditioned Duet Dance GenerationPrerit Gupta, Jason Alexander Fotso-Puepi, Zhengyuan Li, Jay Mehta 等ICCV 2025 · 被引用 1 次
