Cross-domain Knowledge Distillation for Retrieval-based Question Answering Systems
Cen Chen, Chengyu Wang, Minghui Qiu, Dehong Gao, Linbo Jin, Wang Li
Abstract
Question Answering (QA) systems have been extensively studied in both academia and the research community due to their wide real-world applications. When building such industrial-scale QA applications, we are facing two prominent challenges, i.e., i) lacking a sufficient amount of training data to learn an accurate model and ii) requiring high inference speed for online model serving. There are generally two ways to mitigate the above-mentioned problems. One is to adopt transfer learning to leverage information from other domains; the other is to distill the "dark knowledge" from a large teacher model to small student models. The former usually employs parameter sharing mechanisms for knowledge transfer, but does not utilize the "dark knowledge" of pre-trained large models. The latter usually does not consider the cross-domain information from other domains. We argue that these two types of methods can be complementary to each other. Hence in this work, we provide a new perspective on the potential of the teacher-student paradigm facilitating cross-domain transfer learning, where the teacher and student tasks belong to heterogeneous domains, with the goal to improve the student model's performance in the target domain. Our framework considers the "dark knowledge" learned from large teacher models and also leverages the adaptive hints to alleviate the domain differences between teacher and student models. Extensive experiments have been conducted on two text matching tasks for retrieval-based QA systems. Results show the proposed method has better performance than the competing methods including the existing state-of-the-art transfer learning methods. We have also deployed our method in an online production system and observed significant improvements compared to the existing approaches in terms of both accuracy and cross-domain robustness. CCS CONCEPTS • Applied computing → Electronic commerce; • Information systems → Retrieval models and ranking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a23ae69-dac7-4912-9ac2-7b757dcda740Cited by top-tier papers2
- TMLKD: Few-shot Trajectory Metric Learning via Knowledge DistillationDanling Lai, Jiajie Xu, Jianfeng Qu, Pingfu Chao et al.VLDB 2025 · 1 citation
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang et al.ACL 2021
Builds on2
Related papers
- TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence SelectionSiddhant Garg, Thuy Vu, Alessandro MoschittiAAAI 2020 · 229 citations
- DoQA - Accessing Domain-Specific FAQs via Conversational QAJon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu et al.ACL 2020 · 2 citations
- Heterogeneous Continual LearningDivyam Madaan, Hongxu Yin, Wonmin Byeon, Jan Kautz et al.CVPR 2023
- Collaborative Enhancement of Large and Small Models for Question Answering via Dual Knowledge TransferShaofei Wang, Yunan Liu, Xiaolan Tang, Wenlong ChenAAAI 2026
- Learning Systems Expansion with Efficient Heterogeneity-aware Knowledge TransferGaole Dai, Huatao Xu, Yifan Yang, Rui Tan et al.AAAI 2026
