CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video Representations
Mohammadreza Zolfaghari, Yi Zhu, Peter V. Gehler, Thomas Brox
摘要
Contrastive learning allows us to flexibly define powerful losses by contrasting positive pairs from sets of negative samples. Recently, the principle has also been used to learn cross-modal embeddings for video and text, yet without exploiting its full potential. In particular, previous losses do not take the intra-modality similarities into account, which leads to inefficient embeddings, as the same content is mapped to multiple points in the embedding space. With CrossCLR, we present a contrastive loss that fixes this issue. Moreover, we define sets of highly related samples in terms of their input embeddings and exclude them from the negative samples to avoid issues with false negatives. We show that these principles consistently improve the quality of the learned embeddings. The joint embeddings learned with CrossCLR extend the state of the art in video-text retrieval on Youcook2 and LSMDC datasets and in video captioning on Youcook2 dataset by a large margin. We also demonstrate the generality of the concept by learning improved joint embeddings for other pairs of modalities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Decoupled Contrastive Multi-View Clustering with High-Order Random WalksYiding Lu, Yijie Lin, Mouxing Yang, Dezhong Peng 等AAAI 2024 · 被引用 107 次
- AIM: Adapting Image Models for Efficient Video Action RecognitionTaojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang 等ICLR 2023 · 被引用 62 次
- Co-Modality Graph Contrastive Learning for Imbalanced Node ClassificationYiyue Qian, Chunhui Zhang, Yiming Zhang, Qianlong Wen 等NeurIPS 2022 · 被引用 54 次
- Robust Cross-Modal Representation Learning with Progressive Self-DistillationAlex Andonian, Shixing Chen, Raffay HamidCVPR 2022 · 被引用 43 次
- CARAT: Contrastive Feature Reconstruction and Aggregation for Multi-Modal Multi-Label Emotion RecognitionCheng Peng, Ke Chen, Lidan Shou, Gang ChenAAAI 2024 · 被引用 30 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
相关 Paper
- Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited ModalitiesAdriel Saporta, Aahlad Manas Puli, Mark Goldstein, Rajesh RanganathNeurIPS 2024 · 被引用 28 次
- Support-set bottlenecks for video-text representation learningMandela Patrick, Po-Yao Huang, Yuki Markus Asano, Florian Metze 等ICLR 2021 · 被引用 269 次
- Generalized Contrastive Learning for Universal Multimodal RetrievalJungsoo Lee, Janghoon Cho, Hyojin Park, Durga Malladi 等NeurIPS 2025 · 被引用 11 次
- Video Corpus Moment Retrieval with Contrastive LearningHao Zhang, Aixin Sun, Wei Jing, Guoshun Nan 等SIGIR 2021 · 被引用 88 次
- Normalized Contrastive Learning for Text-Video RetrievalYookoon Park, Mahmoud Azab, Seungwhan Moon, Bo Xiong 等EMNLP 2022 · 被引用 10 次
