Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language
Mark Hamilton, Andrew Zisserman, John R. Hershey, William T. Freeman
2024Year
13Top-tier citations
Abstract
Figure 1. Visual overview of the DenseAV algorithm. Two modality-specific backbones featurize audio and visual signals. We introduce a novel generalization of multi-head attention to extract attention maps that discover and separate the "meaning" of spoken words and the sounds an object makes. DenseAV performs this localization and decomposition solely through observing paired stimuli such as videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a019a703-2189-441d-9ba7-0f48c08c81d2Cited by top-tier papers13
- LeGrad: An Explainability Method for Vision Transformers via Feature Formation SensitivityWalid Bousselham, Angie W. Boggust, Sofian Chaybouti, Hendrik Strobelt et al.ICCV 2025 · 47 citations
- Syncphony: Synchronized Audio-to-Video Generation with Diffusion TransformersJibin Song, Mingi Kwon, Jaeseok Jeong, Youngjung UhICLR 2026 · 6 citations
- Can Diffusion Models Disentangle? A Theoretical PerspectiveLiming Wang, Muhammad Jehanzeb Mirza, Yishu Gong, Yuan Gong et al.NeurIPS 2025 · 4 citations
- Seeing Through Touch: Tactile-Driven Visual Localization of Material RegionsSeongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Chung et al.CVPR 2026 · 2 citations
- DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity UnderstandingThomas Kreutz, Max Mühlhäuser, Alejandro Sánchez GuineaICCV 2025 · 1 citation
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
Related papers
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 233 citations
- OV-DAVEL: Towards Open-Vocabulary Dense Audio-Visual Event Localization in Untrimmed VideosJiale Yu, Baopeng Zhang, Zhu Teng, Jianping FanACM MM 2025
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- Watch, Listen and Tell: Multi-Modal Weakly Supervised Dense Event CaptioningTanzila Rahman, Bicheng Xu, Leonid SigalICCV 2019 · 89 citations
- Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and BaselineTiantian Geng, Teng Wang, Jinming Duan, Runmin Cong et al.CVPR 2023
