CAliC: Accurate and Efficient Image-Text Retrieval via Contrastive Alignment and Visual Contexts Modeling
Hongyu Gao, Chao Zhu, Mengyin Liu, Weibo Gu, Hongfa Wang, Wei Liu, Xu-Cheng Yin
摘要
Image-text retrieval is an essential task of information retrieval, in which the models with the Vision-and-Language Pretraining(VLP) are able to achieve ideal accuracy compared with the ones without VLP. Among different VLP approaches, the single-stream models achieve the overall best retrieval accuracy, but slower inference speed. Recently, researchers have introduced the two-stage retrieval setting commonly used in the information retrieval field to the single-stream VLP model for a better accuracy/efficiency trade-off. However, the retrieval accuracy and efficiency are still unsatisfactory mainly due to the limitations of the patch-based visual unimodal encoder in these VLP models. The unimodal encoders are trained on pure visual data, so the visual features extracted by them are difficult to align with the textual features and it is also difficult for the multi-modal encoder to understand visual information. Under these circumstances, we propose an accurate and efficient two-stage image-text retrieval model via Contrastive Alignment and visual Contexts modeling(CAliC). In the first stage of the proposed model, the visual unimodal encoder is pretrained with cross-modal contrastive learning to extract easily aligned visual features, which improves the retrieval accuracy and the inference speed. In the second stage of the proposed model, we introduce a new visual contexts modeling task during pretraining to help the multi-modal encoder better understand the visual information and get more accurate predictions. Extensive experimental evaluation validates the effectiveness of our proposed approach, which achieves a higher retrieval accuracy while keeping a faster inference speed, and outperforms existing state-of-the-art retrieval methods on image-text retrieval tasks over Flickr30K and COCO benchmarks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Tri-Modal Motion Retrieval by Learning a Joint Embedding SpaceKangning Yin, Shihao Zou, Yuxuan Ge, Zheng TianCVPR 2024 · 被引用 7 次
- Seeing What You Miss: Vision-Language Pre-training with Semantic Completion LearningYatai Ji, Rongcheng Tu, Jie Jiang, Weijie Kong 等CVPR 2023
相关 Paper
- COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal RetrievalHaoyu Lu, Nanyi Fei, Yuqi Huo, Yizhao Gao 等CVPR 2022 · 被引用 68 次
- COOKIE: Contrastive Cross-Modal Knowledge Sharing Pre-training for Vision-Language RepresentationKeyu Wen, Jin Xia, Yuanyuan Huang, Linyang Li 等ICCV 2021 · 被引用 35 次
- GilBERT: Generative Vision-Language Pre-Training for Image-Text RetrievalWeixiang Hong, Kaixiang Ji, Jiajia Liu, Jian Wang 等SIGIR 2021 · 被引用 37 次
- Unsupervised Vision-and-Language Pretraining via Retrieval-based Multi-Granular AlignmentMingyang Zhou, Licheng Yu, Amanpreet Singh, Mengjiao Wang 等CVPR 2022 · 被引用 29 次
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 被引用 12 次
