ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning
Sangho Lee, Jiwan Chung, Youngjae Yu, Gunhee Kim, Thomas M. Breuel, Gal Chechik, Yale Song
Abstract
The natural association between visual observations and their corresponding sound provides powerful selfsupervisory signals for learning video representations, which makes the ever-growing amount of online videos an attractive source of training data. However, large portions of online videos contain irrelevant audio-visual signals because of edited/overdubbed audio, and models trained on such uncurated videos have shown to learn suboptimal representations. Therefore, existing approaches rely almost exclusively on datasets with predetermined taxonomies of semantic concepts, where there is a high chance of audiovisual correspondence. Unfortunately, constructing such datasets require labor intensive manual annotation and/or verification, which severely limits the utility of online videos for large-scale learning. In this work, we present an automatic dataset curation approach based on subset optimization where the objective is to maximize the mutual information between audio and visual channels in videos. We demonstrate that our approach finds videos with high audio-visual correspondence and show that self-supervised models trained on our data achieve competitive performances compared to models trained on existing manually curated datasets. The most significant benefit of our approach is scalability: We release ACAV100M that contains 100 million videos with high audio-visual correspondence, ideal for self-supervised video representation learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3ead484-1b21-4433-affc-66a299cd781eCited by top-tier papers19
- Any-to-Any Generation via Composable DiffusionZineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng et al.NeurIPS 2023 · 294 citations
- Robust Contrastive Learning against Noisy ViewsChing-Yao Chuang, R. Devon Hjelm, Xin Wang, Vibhav Vineet et al.CVPR 2022 · 67 citations
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and ActionJiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang et al.CVPR 2024 · 53 citations
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao et al.ICCV 2023 · 51 citations
- Skating-Mixer: Long-Term Sport Audio-Visual Modeling with MLPsJingfei Xia, Mingchen Zhuge, Tiantian Geng, Shun Fan et al.AAAI 2023 · 38 citations
Builds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani et al.NeurIPS 2020 · 483 citations
- Momentum Contrast for Unsupervised Visual Representation LearningKaiming He, Haoqi Fan, Yuxin Wu, Saining Xie et al.CVPR 2020
Related papers
- Telling Left From Right: Learning Spatial Correspondence of Sight and SoundKarren Yang, Bryan C. Russell, Justin SalamonCVPR 2020
- Language-Guided Audio-Visual Source Separation via Trimodal ConsistencyReuben Tan, Arijit Ray, Andrea Burns, Bryan A. Plummer et al.CVPR 2023
- Auto-ACD: A Large-scale Dataset for Audio-Language Representation LearningLuoyi Sun, Xuenan Xu, Mengyue Wu, Weidi XieACM MM 2024 · 23 citations
- Contrastive Audio-Visual Masked AutoencoderYuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath et al.ICLR 2023 · 17 citations
- Enhancing Audio-Visual Association with Self-Supervised Curriculum LearningJingran Zhang, Xing Xu, Fumin Shen, Huimin Lu et al.AAAI 2021 · 22 citations
