PROD: Progressive Distillation for Dense Retrieval
Zhenghao Lin, Yeyun Gong, Xiao Liu, Hang Zhang, Chen Lin, Anlei Dong, Jian Jiao, Jingwen Lu, Daxin Jiang, Rangan Majumder, Nan Duan
Abstract
Knowledge distillation is an effective way to transfer knowledge from a strong teacher to an efficient student model. Ideally, we expect the better the teacher is, the better the student performs. However, this expectation does not always come true. It is common that a strong teacher model results in a bad student via distillation due to the nonnegligible gap between teacher and student. To bridge the gap, we propose PROD, a PROgressive Distillation method, for dense retrieval. PROD consists of a teacher progressive distillation and a data progressive distillation to gradually improve the student. To alleviate catastrophic forgetting, we introduce a regularization term in each distillation process. We conduct extensive experiments on seven datasets including five widely-used publicly available benchmarks: MS MARCO Passage, TREC Passage 19, TREC Document 19, MS MARCO Document, and Natural Questions, as well as two industry datasets: Bing-Rel and Bing-Ads. PROD achieves the state-of-the-art in the distillation methods for dense retrieval. Our 6-layer student model even surpasses most of the existing 12-layer models on all five public benchmarks. The code and models are released in https://github.com/microsoft/SimXNS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20cc8c3b-4296-47d0-a8cf-78acb67a7c8cCited by top-tier papers6
- LED: Lexicon-Enlightened Dense Retriever for Large-Scale RetrievalKai Zhang, Chongyang Tao, Tao Shen, Can Xu et al.WWW 2023 · 27 citations
- Hypencoder: Hypernetworks for Information RetrievalJulian Killingback, Hansi Zeng, Hamed ZamaniSIGIR 2025 · 7 citations
- Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence RegularizationShiqi Wang, Yeqin Zhang, Cam-Tu NguyenAAAI 2024 · 6 citations
- MTA4DPR: Multi-Teaching-Assistants Based Iterative Knowledge Distillation for Dense Passage RetrievalQixi Lu, Endong Xun, Gongbo TangEMNLP 2024 · 2 citations
- Leveraging Estimated Transferability Over Human Intuition for Model Selection in Text RankingJun Bai, Zhuofan Chen, Zhenzi Li, Hanhua Hong et al.EMNLP 2024 · 2 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 337 citations
Related papers
- Empowering Dual-Encoder with Query Generator for Cross-Lingual Dense RetrievalHouxing Ren, Linjun Shou, Ning Wu, Ming Gong et al.EMNLP 2022 · 6 citations
- Bidirectional Distillation for Top-K Recommender SystemWonbin Kweon, SeongKu Kang, Hwanjo YuWWW 2021 · 58 citations
- Channel-wise Knowledge Distillation for Dense Prediction*Changyong Shu, Yifan Liu, Jianfei Gao, Zheng Yan et al.ICCV 2021 · 432 citations
- Refine Myself by Teaching Myself: Feature Refinement via Self-Knowledge DistillationMingi Ji, Seungjae Shin, Seunghyun Hwang, Gibeom Park et al.CVPR 2021
- PairDistill: Pairwise Relevance Distillation for Dense RetrievalChao-Wei Huang, Yun-Nung ChenEMNLP 2024 · 3 citations
