Flipping Knowledge Distillation: Leveraging Small Models' Expertise to Enhance LLMs in Text Matching
Mingzhe Li, Jing Xiang, Qishen Zhang, Kaiyang Wan, Xiuying Chen
Abstract
Knowledge distillation typically involves transferring knowledge from a Large Language Model (LLM) to a Smaller Language Model (SLM). However, in tasks such as text matching, fine-tuned smaller models often yield more effective domain-specific representations, as they focus on optimizing the similarity of input pairs. To leverage both the specialized strengths of small models and the rich semantic understanding of LLMs, we introduce a flipped knowledge distillation paradigm, where LLM learns from SLM. Specifically, we address the architectural gap between decoderonly LLMs and smaller encoder-based models by reinterpreting LLMs in an encoder-decoder manner using LoRA. The encoder generates compressed representations, while the decoder maps them to the output space. During training, the encoder produces representations and their similarities, which are then aligned with the similarity scores produced by the teacher, using our proposed Margin-aware Contrastive Learning (MCL) approach. The MCL ensures accurate similarity for both positive and negative pairs, and adaptively handles the internal differences within positive and negative samples. Our paradigm requires only a reasonably good-performing SLM, allowing the LLM to achieve improved performance. Experiments on financial and healthcare benchmarks, as well as real-world applications, confirm its effectiveness, and the model has been fully deployed in an online environment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2febd4c-0152-43c1-bc62-0d92f61ce53cBuilds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Less is More: Task-aware Layer-wise Distillation for Language Model CompressionChen Liang, Simiao Zuo, Qingru Zhang, Pengcheng He et al.ICML 2023 · 119 citations
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 95 citations
- A Contrastive Framework for Learning Sentence Representations from Pairwise and Triple-wise Perspective in Angular SpaceYuhao Zhang, Hongji Zhu, Yongliang Wang, Nan Xu et al.ACL 2022 · 94 citations
- Modeling User Behavior with Graph Convolution for Personalized Product SearchLu Fan, Qimai Li, Bo Liu, Xiao-Ming Wu et al.WWW 2022 · 24 citations
Related papers
- Test-Time Learning for Large Language ModelsJinwu Hu, Zitian Zhang, Guohao Chen, Xutao Wen et al.ICML 2025
- MLAS-LoRA: Language-Aware Parameters Detection and LoRA-Based Knowledge Transfer for Multilingual Machine TranslationTianyu Dong, Bo Li, Jinsong Liu, Shaolin Zhu et al.ACL 2025
- LoRA-Gen: Specializing Large Language Model via Online LoRA GenerationYicheng Xiao, Lin Song, Rui Yan, Cheng Cheng et al.ICML 2025
- Seeking Neural Nuggets: Knowledge Transfer in Large Language Models from a Parametric PerspectiveMing Zhong, Chenxin An, Weizhu Chen, Jiawei Han et al.ICLR 2024 · 16 citations
- VladVA: Discriminative Fine-tuning of LVLMsYassine Ouali, Adrian Bulat, Alexandros Xenos, Anestis Zaganidis et al.CVPR 2025
