GilBERT: Generative Vision-Language Pre-Training for Image-Text Retrieval
Weixiang Hong, Kaixiang Ji, Jiajia Liu, Jian Wang, Jingdong Chen, Wei Chu
Abstract
Given a text/image query, image-text retrieval aims to find the relevant items in the database. Recently, visual-linguistic pre-training (VLP) methods have demonstrated promising accuracy on image-text retrieval and other visual-linguistic tasks. These VLP methods are typically pre-trained on a large amount of image-text pairs, then fine-tuned on various downstream tasks. Nevertheless, due to the natural modality incompleteness in image-text retrieval, i.e., the query is either image or text rather than an image-text pair, the naive application of VLP to image-text retrieval results in significant inefficiency. Moreover, existing VLP methods cannot extract comparable representations for a single-modal query and multi-modal database items. In this work, we propose a generative visual-linguistic pre-training approach, termed as GilBERT, to simultaneously learn generic representations of image-text data and complete the missing modality for incomplete pairs. In testing phase, the proposed GilBERT facilitates efficient vector-based retrieval by providing unified feature embedding for query and database items. Moreover, the generative training not only makes GilBERT compatible with non-parallel text/image corpus, but also enables GilBERT to model the image-text relationships without suffering massive randomly-sampled negative samples, leading to superior experimental performances. Extensive experiments demonstrate the advantages of GilBERT in image-text retrieval, in terms of both efficiency and accuracy.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers7
- Slide4N: Creating Presentation Slides from Computational Notebooks with Human-AI CollaborationFengjie Wang, Xuye Liu, Oujing Liu, Ali Neshati et al.CHI 2023 · 37 citations
- Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image RetrievalHaokun Wen, Xuemeng Song, Xiaolin Chen, Yinwei Wei et al.SIGIR 2024 · 30 citations
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust LearningYabing Wang, Jianfeng Dong, Tianxiang Liang, Minsong Zhang et al.ACM MM 2022 · 26 citations
- Real20M: A Large-scale E-commerce Dataset for Cross-domain RetrievalYanzhe Chen, Huasong Zhong, Xiangteng He, Yuxin Peng et al.ACM MM 2023 · 15 citations
- CFIR: Fast and Effective Long-Text To Image Retrieval for Large CorporaZijun Long, Xuri Ge, Richard McCreadie, Joemon M. JoseSIGIR 2024 · 10 citations
Related papers
- Unsupervised Vision-and-Language Pretraining via Retrieval-based Multi-Granular AlignmentMingyang Zhou, Licheng Yu, Amanpreet Singh, Mengjiao Wang et al.CVPR 2022 · 29 citations
- CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained KnowledgeLinli Yao, Weijing Chen, Qin JinWWW 2023 · 11 citations
- CAliC: Accurate and Efficient Image-Text Retrieval via Contrastive Alignment and Visual Contexts ModelingHongyu Gao, Chao Zhu, Mengyin Liu, Weibo Gu et al.ACM MM 2022 · 8 citations
- HiVLP: Hierarchical Interactive Video-Language Pre-TrainingBin Shao, Jianzhuang Liu, Renjing Pei, Songcen Xu et al.ICCV 2023 · 6 citations
- Retrieval-based Knowledge Augmented Vision Language Pre-trainingJiahua Rao, Zifei Shan, Longpo Liu, Yao Zhou et al.ACM MM 2023 · 13 citations
