End-to-End Unsupervised Vision-and-Language Pre-training with Referring Expression Matching
Chi Chen, Peng Li, Maosong Sun, Yang Liu
摘要
Recently there has been an emerging interest in unsupervised vision-and-language pre-training (VLP) that learns multimodal representations without parallel image-caption data. These pioneering works significantly reduce the cost of VLP on data collection and achieve promising results compared to supervised VLP. However, existing unsupervised VLP methods take as input pre-extracted region-based visual features from external object detectors, which both limits flexibility and reduces computational efficiency. In this paper, we explore end-to-end unsupervised VLP with a vision encoder to directly encode images. The vision encoder is pre-trained on image-only data and jointly optimized during multimodal pre-training. To further enhance the learned cross-modal features, we propose a novel pre-training task that predicts which patches contain an object referred to in natural language from the encoded visual features. Extensive experiments on four visionand-language tasks show that our approach outperforms previous unsupervised VLP methods and obtains new state-of-the-art results 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Text-based Person Search without Parallel Image-Text DataYang Bai, Jingyao Wang, Min Cao, Chen Chen 等ACM MM 2023 · 被引用 27 次
- Weakly Supervised Vision-and-Language Pre-training with Relative RepresentationsChi Chen, Peng Li, Maosong Sun, Yang LiuACL 2023 · 被引用 3 次
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual LearningHaiyang Xu, Ming Yan, Chenliang Li, Bin Bi 等ACL 2021
- Unsupervised Vision-and-Language Pretraining via Retrieval-based Multi-Granular AlignmentMingyang Zhou, Licheng Yu, Amanpreet Singh, Mengjiao Wang 等CVPR 2022 · 被引用 29 次
- PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language ModelsYuan Yao, Qianyu Chen, Ao Zhang, Wei Ji 等EMNLP 2022 · 被引用 33 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- COPA : Efficient Vision-Language Pre-training through Collaborative Object- and Patch-Text AlignmentChaoya Jiang, Haiyang Xu, Wei Ye, Qinghao Ye 等ACM MM 2023 · 被引用 10 次
