Learning Representations from Audio-Visual Spatial Alignment
Pedro Morgado, Yi Li, Nuno Vasconcelos
摘要
We introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based on audio-visual correspondence (AVC) predict whether audio and video clips originate from the same or different video instances. Audio-visual temporal synchronization (AVTS) further discriminates negative pairs originated from the same video instance but at different moments in time. While these approaches learn high-quality representations for downstream tasks such as action recognition, their training objectives disregard spatial cues naturally occurring in audio and visual signals. To learn from these spatial cues, we tasked a network to perform contrastive audio-visual spatial alignment of 360 video and spatial audio. The ability to perform spatial alignment is enhanced by reasoning over the full spatial content of the 360 video using a transformer architecture to combine representations from multiple viewpoints. The advantages of the proposed pretext task are demonstrated on a variety of audio and visual downstream tasks, including audio-visual correspondence, spatial alignment, action recognition, and video semantic segmentation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video RepresentationsMohammadreza Zolfaghari, Yi Zhu, Peter V. Gehler, Thomas BroxICCV 2021 · 被引用 160 次
- Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video ParsingYan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin 等NeurIPS 2021 · 被引用 94 次
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 被引用 92 次
- Learning Cross-Modal Contrastive Features for Video Domain AdaptationDonghyun Kim, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu 等ICCV 2021 · 被引用 88 次
- SimMMDG: A Simple and Effective Framework for Multi-modal Domain GeneralizationHao Dong, Ismail Nejjar, Han Sun, Eleni N. Chatzi 等NeurIPS 2023 · 被引用 80 次
它引用的顶会 Paper9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani 等NeurIPS 2020 · 被引用 483 次
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- Contrastive Learning with Adversarial ExamplesChih-Hui Ho, Nuno VasconcelosNeurIPS 2020 · 被引用 174 次
相关 Paper
- Bio-Inspired Audiovisual Multi-Representation Integration via Self-Supervised LearningZhaojian Li, Bin Zhao, Yuan YuanACM MM 2023 · 被引用 3 次
- Telling Left From Right: Learning Spatial Correspondence of Sight and SoundKarren Yang, Bryan C. Russell, Justin SalamonCVPR 2020
- Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningYing Cheng, Ruize Wang, Zhihao Pan, Rui Feng 等ACM MM 2020 · 被引用 93 次
- Audio-Visual Instance Discrimination with Cross-Modal AgreementPedro Morgado, Nuno Vasconcelos, Ishan MisraCVPR 2021
- Spatial-then-Temporal Self-Supervised Learning for Video CorrespondenceRui Li, Dong LiuCVPR 2023
