Unsupervised Vision-and-Language Pretraining via Retrieval-based Multi-Granular Alignment
Mingyang Zhou, Licheng Yu, Amanpreet Singh, Mengjiao Wang, Zhou Yu, Ning Zhang
摘要
Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require pre-training on a large set of parallel imagetext data, which is costly to collect, compared to image-only or text-only data. In this paper, we explore unsupervised Vision-and-Language pre-training (UVLP) to learn the cross-modal representation from non-parallel image and text datasets. We found two key factors that lead to good unsupervised V + L pre-training without parallel data: (i) joint image-and-text input (ii) overall imagetext alignment (even for non-parallel data). Accordingly, we propose a novel unsupervised V + L pre-training curriculum for non-parallel texts and images. We first construct a weakly aligned imagetext corpus via a retrieval-based approach, then apply a set of multi-granular alignment pre-training tasks, including region-to-tag, region-to-phrase, and image-to-sentence alignment, to bridge the gap between the two modalities. A comprehensive ablation study shows each granularity is helpful to learn a stronger pre-trained model. We adapt our pre-trained model to a set of V+L downstream tasks, including VQA, NLVR2, Visual Entailment, and Ref-COCO+. Our model achieves the state-of-art performance in all these tasks under the unsupervised setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- S-CLIP: Semi-supervised Vision-Language Learning using Few Specialist CaptionsSangwoo Mo, Minkyu Kim, Kyungmin Lee, Jinwoo ShinNeurIPS 2023 · 被引用 53 次
- Bootstrapping Vision-Language Learning with Decoupled Language Pre-trainingYiren Jian, Chongyang Gao, Soroush VosoughiNeurIPS 2023 · 被引用 48 次
- Text-based Person Search without Parallel Image-Text DataYang Bai, Jingyao Wang, Min Cao, Chen Chen 等ACM MM 2023 · 被引用 27 次
- End-to-End Unsupervised Vision-and-Language Pre-training with Referring Expression MatchingChi Chen, Peng Li, Maosong Sun, Yang LiuEMNLP 2022 · 被引用 7 次
- Weakly Supervised Vision-and-Language Pre-training with Relative RepresentationsChi Chen, Peng Li, Maosong Sun, Yang LiuACL 2023 · 被引用 3 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- Unifying Cross-Lingual and Cross-Modal Modeling Towards Weakly Supervised Multilingual Vision-Language Pre-trainingZejun Li, Zhihao Fan, Jingjing Chen, Qi Zhang 等ACL 2023 · 被引用 12 次
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual LearningHaiyang Xu, Ming Yan, Chenliang Li, Bin Bi 等ACL 2021
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- M2-VLP: Enhancing Multilingual Vision-Language Pre-Training via Multi-Grained AlignmentAhtamjan Ahmat, Lei Wang, Yating Yang, Bo Ma 等WWW 2025 · 被引用 2 次
- Seeing What You Miss: Vision-Language Pre-training with Semantic Completion LearningYatai Ji, Rongcheng Tu, Jie Jiang, Weijie Kong 等CVPR 2023
