Let All Be Whitened: Multi-Teacher Distillation for Efficient Visual Retrieval
Zhe Ma, Jianfeng Dong, Shouling Ji, Zhenguang Liu, Xuhong Zhang, Zonghui Wang, Sifeng He, Feng Qian, Xiaobo Zhang, Lei Yang
Abstract
Visual retrieval aims to search for the most relevant visual items, e.g., images and videos, from a candidate gallery with a given query item. Accuracy and efficiency are two competing objectives in retrieval tasks. Instead of crafting a new method pursuing further improvement on accuracy, in this paper we propose a multi-teacher distillation framework Whiten-MTD, which is able to transfer knowledge from off-the-shelf pre-trained retrieval models to a lightweight student model for efficient visual retrieval. Furthermore, we discover that the similarities obtained by different retrieval models are diversified and incommensurable, which makes it challenging to jointly distill knowledge from multiple models. Therefore, we propose to whiten the output of teacher models before fusion, which enables effective multi-teacher distillation for retrieval models. Whiten-MTD is conceptually simple and practically effective. Extensive experiments on two landmark image retrieval datasets and one video retrieval dataset demonstrate the effectiveness of our proposed method, and its good balance of retrieval performance and efficiency. Our source code is released at https://github.com/Maryeon/whiten_mtd.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4af7bced-2ec6-4439-aca9-5fd0bc7fbfc3Cited by top-tier papers6
- UDON: Universal Dynamic Online distillatioN for generic image representationsNikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araújo, Ondrej ChumNeurIPS 2024 · 12 citations
- TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent CollaborationYiwei Guo, Shaobin Zhuang, Kunchang Li, Yu Qiao et al.NeurIPS 2024 · 9 citations
- Learning Task-Agnostic Representations through Multi-Teacher DistillationPhilippe Formont, Maxime Darrin, Banafsheh Karimian, Eric Granger et al.NeurIPS 2025 · 6 citations
- LLM-Assisted Entropy-Based Adaptive Distillation for Unsupervised Fine-Grained Visual Representation LearningJianfeng Dong, Danfeng Luo, Daizong Liu, Jie Sun et al.ICCV 2025 · 1 citation
- Single Teacher, Multiple Perspectives: Teacher Knowledge Augmentation for Enhanced Knowledge DistillationMd. Imtiaz Hossain, Sharmen Akhter, Choong Seon Hong, Eui-Nam HuhICLR 2025
Builds on29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- MTA4DPR: Multi-Teaching-Assistants Based Iterative Knowledge Distillation for Dense Passage RetrievalQixi Lu, Endong Xun, Gongbo TangEMNLP 2024 · 2 citations
- Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual RecognitionChuanguang Yang, Xinqiang Yu, Han Yang, Zhulin An et al.AAAI 2025 · 26 citations
- Dual-Teacher Interactive Knowledge Distillation Network for Text-to-Visible & Infrared Person RetrievalChenglong Li, Zhengyu Chen, Yifei Deng, Aihua ZhengAAAI 2026
- Multi-Mode Online Knowledge Distillation for Self-Supervised Visual Representation LearningKaiyou Song, Jin Xie, Shan Zhang, Zimeng LuoCVPR 2023
- Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video RetrievalJianfeng Dong, Minsong Zhang, Zheng Zhang, Xianke Chen et al.ICCV 2023 · 35 citations
