Audiovisual Masked Autoencoders
Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, Anurag Arnab
摘要
Can we leverage the audiovisual information already present in video to improve self-supervised representation learning? To answer this question, we study various pretraining architectures and objectives within the masked autoencoding framework, motivated by the success of similar methods in natural language and image understanding. We show that we can achieve significant improvements on audiovisual downstream classification tasks, surpassing the state-of-the-art on VGGSound and AudioSet. Furthermore, we can leverage our audiovisual pretraining scheme for multiple unimodal downstream tasks using a single audiovisual pretrained model. We additionally demonstrate the transferability of our representations, achieving state-of-the-art audiovisual results on Epic Kitchens without pretraining specifically for this dataset. To facilitate further research, we have released code and models at https://github.com/google-research/scenic .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- AVFF: Audio-Visual Feature Fusion for Video Deepfake DetectionTrevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki 等CVPR 2024 · 被引用 51 次
- A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual GenerationGwanghyun Kim, Alonso Martinez, Yu-Chuan Su, Brendan Jou 等NeurIPS 2024 · 被引用 23 次
- From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and GenerationKun Su, Xiulong Liu, Eli ShlizermanICML 2024 · 被引用 22 次
- Unveiling the Power of Audio-Visual Early Fusion Transformers with Dense Interactions Through Masked ModelingShentong Mo, Pedro MorgadoCVPR 2024 · 被引用 20 次
- EquiAV: Leveraging Equivariance for Audio-Visual Contrastive LearningJongsuk Kim, Hyeongkeun Lee, Kyeongha Rho, Junmo Kim 等ICML 2024 · 被引用 15 次
它引用的顶会 Paper38
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
相关 Paper
- Contrastive Audio-Visual Masked AutoencoderYuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath 等ICLR 2023 · 被引用 17 次
- MAViL: Masked Audio-Video LearnersPo-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya Ryali 等NeurIPS 2023 · 被引用 95 次
- Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningYing Cheng, Ruize Wang, Zhihao Pan, Rui Feng 等ACM MM 2020 · 被引用 93 次
- Learning State-Aware Visual Representations from Audible InteractionsHimangi Mittal, Pedro Morgado, Unnat Jain, Abhinav GuptaNeurIPS 2022 · 被引用 30 次
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski 等NeurIPS 2022 · 被引用 524 次
