PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano Performance
Qijun Gan, Song Wang, Shengtao Wu, Jianke Zhu
摘要
Recently, artificial intelligence techniques for education have been received increasing attentions, while it still remains an open problem to design the effective music instrument instructing systems. Although key presses can be directly derived from sheet music, the transitional movements among key presses require more extensive guidance in piano performance. In this work, we construct a piano-hand motion generation benchmark to guide hand movements and fingerings for piano playing. To this end, we collect an annotated dataset, PianoMotion10M, consisting of 116 hours of piano playing videos from a bird's-eye view with 10 million annotated hand poses. We also introduce a powerful baseline model that generates hand motions from piano audios through a position predictor and a position-guided gesture generator. Furthermore, a series of evaluation metrics are designed to assess the performance of the baseline model, including motion similarity, smoothness, positional accuracy of left and right hands, and overall fidelity of movement distribution. Despite that piano key presses with respect to music scores or audios are already accessible, PianoMotion10M aims to provide guidance on piano fingering for instruction purposes. The source code and dataset can be accessed at https://github.com/agnJason/PianoMotion10M.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- From Pose to Muscle: Multimodal Learning for Piano Hand Muscle ElectromyographyRuofan Liu, Yichen Peng, Takanori Oku, Chen-Chieh Liao 等NeurIPS 2025 · 被引用 6 次
- Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion SynthesisZihao Liu, Mingwen Ou, Zunnan Xu, Jiaqi Huang 等ACM MM 2025 · 被引用 2 次
- MUSIC: Learning Muscle-Driven Dexterous Hand ControlPei Xu, Yufei Ye, Shuchun Sun, Yu Ding 等SIGGRAPH 2026
它引用的顶会 Paper30
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 被引用 701 次
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 被引用 687 次
相关 Paper
- Automatic Piano Fingering from Partially Annotated Scores using Autoregressive Neural NetworksPedro Ramoneda, Dasaem Jeong, Eita Nakamura, Xavier Serra 等ACM MM 2022 · 被引用 7 次
- Aria-MIDI: A Dataset of Piano MIDI Files for Symbolic Music ModelingLouis Bradshaw, Simon ColtonICLR 2025
- Audeo: Audio Generation for a Silent Performance VideoKun Su, Xiulong Liu, Eli ShlizermanNeurIPS 2020 · 被引用 78 次
- Visualising Pianists' Touch: Transcribing Expressive Piano Performance from Audio to Piano Key MotionJingjing Tang, Shinichi Furuya, Hayato Nishioka, Momoko Shioki 等CHI 2026 · 被引用 1 次
- Audio Matters Too! Enhancing Markerless Motion Capture with Audio Signals for String Performance CaptureYitong Jin, Zhiping Qiu, Yi Shi, Shuangpeng Sun 等SIGGRAPH 2024 · 被引用 9 次
