Active Contrastive Learning of Audio-Visual Video Representations
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale Song
Abstract
Contrastive learning has been shown to produce generalizable representations of audio and visual data by maximizing the lower bound on the mutual information (MI) between different views of an instance. However, obtaining a tight lower bound requires a sample size exponential in MI and thus a large set of negative samples. We can incorporate more samples by building a large queue-based dictionary, but there are theoretical limits to performance improvements even with a large number of negative samples. We hypothesize that random negative sampling leads to a highly redundant dictionary that results in suboptimal representations for downstream tasks. In this paper, we propose an active contrastive learning approach that builds an actively sampled dictionary with diverse and informative items, which improves the quality of negative samples and improves performances on tasks where there is high mutual information in the data, e.g., video classification. Our model achieves state-of-the-art performance on challenging audio and visual downstream benchmarks including UCF101, HMDB51 and ESC50. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers44
- Leveraging Real Talking Faces via Self-Supervision for Robust Forgery DetectionAlexandros Haliassos, Rodrigo Mira, Stavros Petridis, Maja PanticCVPR 2022 · 138 citations
- MAViL: Masked Audio-Video LearnersPo-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya Ryali et al.NeurIPS 2023 · 95 citations
- Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video ParsingYan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin et al.NeurIPS 2021 · 94 citations
- COCOA: Cross Modality Contrastive Learning for Sensor DataShohreh Deldari, Hao Xue, Aaqib Saeed, Daniel V. Smith et al.UbiComp 2022 · 88 citations
- On Learning Contrastive Representations for Learning with Noisy LabelsLi Yi, Sheng Liu, Qi She, A. Ian McLeod et al.CVPR 2022 · 68 citations
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
Related papers
- Audio-Visual Instance Discrimination with Cross-Modal AgreementPedro Morgado, Nuno Vasconcelos, Ishan MisraCVPR 2021
- Contrastive Learning of Global and Local Video RepresentationsShuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale SongNeurIPS 2021 · 45 citations
- Unraveling Instance Associations: A Closer Look for Audio-Visual SegmentationYuanhong Chen, Yuyuan Liu, Hu Wang, Fengbei Liu et al.CVPR 2024
- On the Surrogate Gap between Contrastive and Supervised LossesHan Bao, Yoshihiro Nagano, Kento NozawaICML 2022 · 27 citations
- Conditional Negative Sampling for Contrastive Learning of Visual RepresentationsMike Wu, Milan Mossé, Chengxu Zhuang, Daniel Yamins et al.ICLR 2021 · 89 citations
