Music Gesture for Visual Sound Separation
Chuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum, Antonio Torralba
Abstract
Recent deep learning approaches have achieved impressive performance on visual sound separation tasks. However, these approaches are mostly built on appearance and optical flow like motion feature representations, which exhibit limited abilities to find the correlations between audio signals and visual points, especially when separating multiple instruments of the same types, such as multiple violins in a scene. To address this, we propose ``Music Gesture," a keypoint-based structured representation to explicitly model the body and finger movements of musicians when they perform music. We first adopt a context-aware graph network to integrate visual semantic context with body dynamics and then apply an audio-visual fusion model to associate body movements with the corresponding audio signals. Experimental results on three music performance datasets show: 1) strong improvements upon benchmark metrics for hetero-musical separation tasks (i.e. different instruments); 2) new ability for effective homo-musical separation for piano, flute, and trumpet duets, which to our best knowledge has never been achieved with alternative methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0cb52607-7006-47ca-b194-69285a8ea11bCited by top-tier papers61
- Learning Representations from Audio-Visual Spatial AlignmentPedro Morgado, Yi Li, Nuno VasconcelosNeurIPS 2020 · 149 citations
- Active Contrastive Learning of Audio-Visual Video RepresentationsShuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale SongICLR 2021 · 109 citations
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
- Cross-Modal Relation-Aware Networks for Audio-Visual Event LocalizationHaoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan et al.ACM MM 2020 · 97 citations
- Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video ParsingYan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin et al.NeurIPS 2021 · 94 citations
Builds on4
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
- Recursive Visual Sound Separation Using Minus-Plus NetXudong Xu, Bo Dai, Dahua LinICCV 2019 · 95 citations
Related papers
- Visual Scene Graphs for Audio Source SeparationMoitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja, Anoop CherianICCV 2021 · 45 citations
- Learning Audio-Visual Dynamics Using Scene Graphs for Audio Source SeparationMoitreya Chatterjee, Narendra Ahuja, Anoop CherianNeurIPS 2022 · 16 citations
- Emphasizing Semantic Consistency of Salient Posture for Speech-Driven Gesture GenerationFengqi Liu, Hexiang Wang, Jingyu Gong, Ran Yi et al.ACM MM 2024 · 2 citations
- TriBERT: Human-centric Audio-visual Representation LearningTanzila Rahman, Mengyu Yang, Leonid SigalNeurIPS 2021 · 15 citations
- Vision-guided Music Source Separation via a Fine-grained Cycle-Separation NetworkShuo Ma, Yanli Ji, Xing Xu, Xiaofeng ZhuACM MM 2021 · 4 citations
