How Much Can CLIP Benefit Vision-and-Language Tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, Kurt Keutzer
摘要
Most existing Vision-and-Language (V&L) models rely on pre-trained visual encoders, using a relatively small set of manually-annotated data (as compared to web-crawled data), to perceive the visual world. However, it has been observed that large-scale pretraining usually can result in better generalization performance, e.g., CLIP (Contrastive Language-Image Pre-training), trained on a massive amount of image-caption pairs, has shown a strong zero-shot capability on various vision tasks. To further study the advantage brought by CLIP, we propose to use CLIP as the visual encoder in various V&L models in two typical scenarios: 1) plugging CLIP into task-specific fine-tuning; 2) combining CLIP with V&L pre-training and transferring to downstream tasks. We show that CLIP significantly outperforms widely-used visual encoders trained with in-domain annotated data, such as BottomUp-TopDown. We achieve competitive or better results on diverse V&L tasks, while establishing new state-of-the-art results on Visual Question Answering, Visual Entailment, and V&L Navigation tasks. We release our code at https://github.com/clip-vil/CLIP-ViL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper92
- Patching open-vocabulary models by interpolating weightsGabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song 等NeurIPS 2022 · 被引用 230 次
- Coarse-to-Fine Vision-Language Pre-training with Fusion in the BackboneZi-Yi Dou, Aishwarya Kamath, Zhe Gan, Pengchuan Zhang 等NeurIPS 2022 · 被引用 173 次
- ReCLIP: A Strong Zero-Shot Baseline for Referring Expression ComprehensionSanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner 等ACL 2022 · 被引用 172 次
- Open-Vocabulary Universal Image Segmentation with MaskCLIPZheng Ding, Jieke Wang, Zhuowen TuICML 2023 · 被引用 150 次
- CenterCLIP: Token Clustering for Efficient Text-Video RetrievalShuai Zhao, Linchao Zhu, Xiaohan Wang, Yi YangSIGIR 2022 · 被引用 150 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
相关 Paper
- CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual EntailmentHaoyu Song, Li Dong, Weinan Zhang, Ting Liu 等ACL 2022
- RWKV-CLIP: A Robust Vision-Language Representation LearnerTiancheng Gu, Kaicheng Yang, Xiang An, Ziyong Feng 等EMNLP 2024 · 被引用 11 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
- LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality RepresentationWeiquan Huang, Aoqi Wu, Yifan Yang, Xufang Luo 等AAAI 2026
- Modeling Caption Diversity in Contrastive Vision-Language PretrainingSamuel Lavoie, Polina Kirichenko, Mark Ibrahim, Mido Assran 等ICML 2024 · 被引用 44 次
