Contrastive Learning of Global and Local Video Representations
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale Song
摘要
Contrastive learning has delivered impressive results for various tasks in the selfsupervised regime. However, existing approaches optimize for learning representations specific to downstream scenarios, i.e., global representations suitable for tasks such as classification or local representations for tasks such as detection and localization. While they produce satisfactory results in the intended downstream scenarios, they often fail to generalize to tasks that they were not originally designed for. In this work, we propose to learn video representations that generalize to both the tasks which require global semantic information (e.g., classification) and the tasks that require local fine-grained spatio-temporal information (e.g., localization). We achieve this by optimizing two contrastive objectives that together encourage our model to learn global-local visual information given audio signals. We show that the two objectives mutually improve the generalizability of the learned global-local representations, significantly outperforming their disjointly learned counterparts. We demonstrate our approach on various tasks including action/sound classification, lip reading, deepfake detection, event and sound localization. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- MAViL: Masked Audio-Video LearnersPo-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya Ryali 等NeurIPS 2023 · 被引用 95 次
- Cross-modal Prompts: Adapting Large Pre-trained Models for Audio-Visual Downstream TasksHaoyi Duan, Yan Xia, Mingze Zhou, Li Tang 等NeurIPS 2023 · 被引用 59 次
- Finding Order in Chaos: A Novel Data Augmentation Method for Time Series in Contrastive LearningBerken Utku Demirel, Christian HolzNeurIPS 2023 · 被引用 48 次
- Learning State-Aware Visual Representations from Audible InteractionsHimangi Mittal, Pedro Morgado, Unnat Jain, Abhinav GuptaNeurIPS 2022 · 被引用 30 次
- SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery DetectionYachao Liang, Min Yu, Gang Li, Jianguo Jiang 等NeurIPS 2024 · 被引用 19 次
它引用的顶会 Paper31
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
相关 Paper
- Audio-Visual Instance Discrimination with Cross-Modal AgreementPedro Morgado, Nuno Vasconcelos, Ishan MisraCVPR 2021
- Audio-Visual Contrastive Learning with Temporal Self-SupervisionSimon Jenni, Alexander Black, John P. CollomosseAAAI 2023 · 被引用 25 次
- Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationPeijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er 等AAAI 2023 · 被引用 13 次
- CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained AlignmentEdson Araujo, Andrew Rouditchenko, Yuan Gong, Saurabhchand Bhati 等CVPR 2025
- Contextualized Spatio-Temporal Contrastive Learning with Self-SupervisionLiangzhe Yuan, Rui Qian, Yin Cui, Boqing Gong 等CVPR 2022 · 被引用 24 次
