Co-Separating Sounds of Visual Objects
Ruohan Gao, Kristen Grauman
Abstract
Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidestep the issue by training with artificially mixed video clips, but this puts unwieldy restrictions on training data collection and may even prevent learning the properties of "true" mixed sounds. We introduce a co-separation training paradigm that permits learning object-level sounds from unlabeled multi-source videos. Our novel training objective requires that the deep neural network's separated audio for similar-looking objects be consistently identifiable, while simultaneously reproducing accurate video-level audio tracks for each source training pair. Our approach disentangles sounds in realistic test videos, even in cases where an object was not observed individually during training. We obtain state-of-the-art results on visually-guided audio source separation and audio denoising for the MUSIC, AudioSet, and AV-Bench datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff6bafd1-7d85-4358-994c-398c3ff285fcCited by top-tier papers62
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Unsupervised Sound Separation Using Mixture Invariant TrainingScott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss et al.NeurIPS 2020 · 227 citations
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan et al.NeurIPS 2020 · 156 citations
- Learning Representations from Audio-Visual Spatial AlignmentPedro Morgado, Yi Li, Nuno VasconcelosNeurIPS 2020 · 149 citations
- Active Contrastive Learning of Audio-Visual Video RepresentationsShuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale SongICLR 2021 · 109 citations
Builds on1
Related papers
- Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen SoundsEfthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey et al.ICLR 2021 · 83 citations
- Visual Scene Graphs for Audio Source SeparationMoitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja, Anoop CherianICCV 2021 · 45 citations
- Audio-Visual Grouping Network for Sound Localization from MixturesShentong Mo, Yapeng TianCVPR 2023
- Language-Guided Audio-Visual Source Separation via Trimodal ConsistencyReuben Tan, Arijit Ray, Andrea Burns, Bryan A. Plummer et al.CVPR 2023
- Cyclic Co-Learning of Sounding Object Visual Grounding and Sound SeparationYapeng Tian, Di Hu, Chenliang XuCVPR 2021
