Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenes
Hyeonggon Ryu, Seongyu Kim, Joon Son Chung, Arda Senocak
摘要
We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited to handling either speech or non-speech sounds independently, or at best, together but sequentially without mixing. This limitation prevents them from capturing the complexity of real-world audio sources that are often mixed. Our approach introduces a "mix-and-separate" framework with audio-visual alignment objectives that jointly learn correspondence and disentanglement using mixed audio. Through these objectives, our model learns to produce distinct embeddings for each audio type, enabling effective disentanglement and grounding across mixed audio sources. Additionally, we created a new dataset to evaluate simultaneous grounding of mixed audio sources, demonstrating that our model outperforms prior methods. Our approach also achieves comparable or better performance in standard segmentation and cross-modal retrieval tasks, highlighting the benefits of our mix-and-separate approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Hear you are: Teaching LLMs Spatial Reasoning with Vision and Spatial SoundHyeonggon Ryu, Joon Son Chung, David HarwathCVPR 2026 · 被引用 4 次
- Seeing Through Touch: Tactile-Driven Visual Localization of Material RegionsSeongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Chung 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper22
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan 等NeurIPS 2020 · 被引用 156 次
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 被引用 92 次
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded SpeechDavid Harwath, Wei-Ning Hsu, James R. GlassICLR 2020 · 被引用 88 次
相关 Paper
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- Cross-Modal Mutual Learning for Audio-Visual Speech Recognition and ManipulationChih-Chun Yang, Wan-Cyuan Fan, Cheng-Fu Yang, Yu-Chiang Frank WangAAAI 2022 · 被引用 16 次
- Sound Source Localization is All about Cross-Modal AlignmentArda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh 等ICCV 2023 · 被引用 39 次
- OmniSep: Unified Omni-Modality Sound Separation with Query-MixupXize Cheng, Siqi Zheng, Zehan Wang, Minghui Fang 等ICLR 2025
