The Sound of Motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio Torralba
摘要
Sounds originate from object motions and vibrations of surrounding air. Inspired by the fact that humans is capable of interpreting sound sources from how objects move visually, we propose a novel system that explicitly captures such motion cues for the task of sound localization and separation. Our system is composed of an end-to-end learnable model called Deep Dense Trajectory (DDT), and a curriculum learning scheme. It exploits the inherent coherence of audio-visual signals from a large quantities of unlabeled videos. Quantitative and qualitative evaluations show that comparing to previous models that rely on visual appearance cues, our motion based system improves performance in separating musical instrument sounds. Furthermore, it separates sound components from duets of the same category of instruments, a challenging problem that has not been addressed before.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper79
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- Labelling unlabelled videos from scratch with multi-modal self-supervisionYuki Markus Asano, Mandela Patrick, Christian Rupprecht, Andrea VedaldiNeurIPS 2020 · 被引用 169 次
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang 等ICML 2022 · 被引用 168 次
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox 等ICCV 2019 · 被引用 157 次
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan 等NeurIPS 2020 · 被引用 156 次
相关 Paper
- Music Gesture for Visual Sound SeparationChuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum 等CVPR 2020
- A Unified Audio-Visual Learning Framework for Localization, Separation, and RecognitionShentong Mo, Pedro MorgadoICML 2023 · 被引用 27 次
- Language-Guided Audio-Visual Source Separation via Trimodal ConsistencyReuben Tan, Arijit Ray, Andrea Burns, Bryan A. Plummer 等CVPR 2023
- Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and LanguageMark Hamilton, Andrew Zisserman, John R. Hershey, William T. FreemanCVPR 2024
- Visual Scene Graphs for Audio Source SeparationMoitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja, Anoop CherianICCV 2021 · 被引用 45 次
