Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
Zihao Liu, Mingwen Ou, Zunnan Xu, Jiaqi Huang, Haonan Han, Ronghui Li, Xiu Li
Abstract
Automating the synthesis of coordinated bimanual piano performances poses significant challenges, particularly in capturing the intricate choreography between the hands while preserving their distinct kinematic signatures. In this paper, we propose a dual-stream neural framework designed to generate synchronized hand gestures for piano playing from audio input, addressing the critical challenge of modeling both hand independence and coordination. Our framework introduces two key innovations: (i) a decoupled diffusion-based generation framework that independently models each hand's motion via dual-noise initialization, sampling distinct latent noise for each while leveraging a shared positional condition, and (ii) a Hand-Coordinated Asymmetric Attention (HCAA) mechanism suppresses symmetric (common-mode) noise to highlight asymmetric handspecific features, while adaptively enhancing inter-hand coordination during denoising. Comprehensive evaluations demonstrate that our framework outperforms existing stateof-the-art methods across multiple metrics. Our project is available at S2C.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction GraphsBin Li, Ruichi Zhang, Han Liang, Jingyan Zhang et al.CVPR 2026 · 4 citations
- MUSIC: Learning Muscle-Driven Dexterous Hand ControlPei Xu, Yufei Ye, Shuchun Sun, Yu Ding et al.SIGGRAPH 2026
Builds on40
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- PhysDiff: Physics-Guided Human Motion Diffusion ModelYe Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat et al.ICCV 2023 · 414 citations
Related papers
- Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic ConditioningHao Zheng, Hu Wang, Tiantian Zheng, Prajjwal Bhattarai et al.CVPR 2026 · 2 citations
- DialoGen: Towards Dialog Gesture Generation via Identity-Decoupled Style Guidance in Interactive Diffusion ModelWeiyu Zhao, Chenyang Wang, Liangxiao Hu, Zonglin Li et al.AAAI 2026
- Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion ModelXu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin et al.CVPR 2024
- GestureHYDRA: Semantic Co-Speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented GenerationQuanwei Yang, Luying Huang, Kaisiyuan Wang, Jiazhi Guan et al.ICCV 2025 · 5 citations
- DIDiffGes: Decoupled Semi-Implicit Diffusion Models for Real-time Gesture Generation from SpeechYongkang Cheng, Shaoli Huang, Xuelin Chen, Jifeng Ning et al.AAAI 2025 · 3 citations
