How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?
Yuxin Chen, Zongyang Ma, Ziqi Zhang, Zhongang Qi, Chunfeng Yuan, Bing Li, Junfu Pu, Ying Shan, Xiaojuan Qi, Weiming Hu
Abstract
Dominant dual-encoder models enable efficient imagetext retrieval but suffer from limited accuracy, while the cross-encoder models offer higher accuracy at the expense of efficiency. Distilling cross-modality matching knowledge from cross-encoder to dual-encoder provides a natural approach to harness their strengths. Thus, we investigate the following valuable question: how to make crossencoder a good teacher for dual-encoder? Our findings are threefold: (1) Cross-modal similarity score distribution of cross-encoder is more concentrated, while the result of dual-encoder is nearly normal, making vanilla logit distillation less effective. However, ranking distillation remains practical, as it is not affected by the score distribution. (2) Only the relative order between hard negatives conveys valid knowledge, while the order information between easy negatives has little significance. (3) Maintaining the coordination between distillation loss and dual-encoder training loss is beneficial for knowledge transfer. Based on these findings, we propose a novel Contrastive Partial Ranking Distillation (CPRD) method, which implements the objective of mimicking relative order between hard negative samples with contrastive learning. This approach coordinates with the training of the dual-encoder, effectively transferring valid knowledge from the cross-encoder to the dualencoder. Extensive experiments on image-text retrieval and ranking tasks show that our method surpasses other distillation methods and significantly improves the accuracy of dual-encoder.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext edaca9d7-a237-47aa-8194-9b97bbcb9412Cited by top-tier papers2
- Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement LearningHaonan Jia, Shichao Dong, Xin Dong, Zenghui Sun et al.CVPR 2026
- Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level CaptionsChan Hur, Jeong-Hun Hong, Dong-hun Lee, Dabin Kang et al.CVPR 2025
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
Related papers
- Complementary Relation Contrastive DistillationJinguo Zhu, Shixiang Tang, Dapeng Chen, Shijie Yu et al.CVPR 2021
- PLD: A Choice-Theoretic List-Wise Knowledge DistillationEjafa Bassam, Dawei Zhu, Kaigui BianNeurIPS 2025
- Wasserstein Contrastive Representation DistillationLiqun Chen, Dong Wang, Zhe Gan, Jingjing Liu et al.CVPR 2021
- RankCSE: Unsupervised Sentence Representations Learning via Learning to RankJiduan Liu, Jiahao Liu, Qifan Wang, Jingang Wang et al.ACL 2023 · 30 citations
- Semantic Distillation from Neighborhood for Composed Image RetrievalYifan Wang, Wuliang Huang, Lei Li, Chun YuanACM MM 2024 · 7 citations
