Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation Learning
Ying Cheng, Ruize Wang, Zhihao Pan, Rui Feng, Yuejie Zhang
摘要
When watching videos, the occurrence of a visual event is often accompanied by an audio event, e.g., the voice of lip motion, the music of playing instruments. There is an underlying correlation between audio and visual events, which can be utilized as free supervised information to train a neural network by solving the pretext task of audio-visual synchronization. In this paper, we propose a novel self-supervised framework with co-attention mechanism to learn generic cross-modal representations from unlabelled videos in the wild, and further benefit downstream tasks. Specifically, we explore three different co-attention modules to focus on discriminative visual regions correlated to the sounds and introduce the interactions between them. Experiments show that our model achieves state-of-the-art performance on the pretext task while having fewer parameters compared with existing methods. To further evaluate the generalizability and transferability of our approach, we apply the pre-trained model on two downstream tasks, i.e., sound source localization and action recognition. Extensive experiments demonstrate that our model provides competitive results with other self-supervised methods, and also indicate that our approach can tackle the challenging scenes which contain multiple sound sources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Hard Negative Mixing for Contrastive LearningYannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel 等NeurIPS 2020 · 被引用 805 次
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 等ACM MM 2021 · 被引用 154 次
- SimMMDG: A Simple and Effective Framework for Multi-modal Domain GeneralizationHao Dong, Ismail Nejjar, Han Sun, Eleni N. Chatzi 等NeurIPS 2023 · 被引用 80 次
- Modality-aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence DetectionJiashuo Yu, Jinyu Liu, Ying Cheng, Rui Feng 等ACM MM 2022 · 被引用 64 次
- MM-Pyramid: Multimodal Pyramid Attentional Network for Audio-Visual Event Localization and Video ParsingJiashuo Yu, Ying Cheng, Rui-Wei Zhao, Rui Feng 等ACM MM 2022 · 被引用 62 次
它引用的顶会 Paper1
相关 Paper
- Learning Representations from Audio-Visual Spatial AlignmentPedro Morgado, Yi Li, Nuno VasconcelosNeurIPS 2020 · 被引用 149 次
- Bio-Inspired Audiovisual Multi-Representation Integration via Self-Supervised LearningZhaojian Li, Bin Zhao, Yuan YuanACM MM 2023 · 被引用 3 次
- Audiovisual Masked AutoencodersMariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic 等ICCV 2023 · 被引用 60 次
- Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationPeijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er 等AAAI 2023 · 被引用 13 次
- Cross-modal Background Suppression for Audio-Visual Event LocalizationYan Xia, Zhou ZhaoCVPR 2022 · 被引用 62 次
