Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation
Jiachen Zhao, Wenlong Zhao, Andrew Drozdov, Benjamin Rozonoyer, Md. Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum
Abstract
We study semi-supervised sequence generation tasks, where the few labeled examples are too scarce to finetune a model, and meanwhile, fewshot prompted large language models (LLMs) exhibit room for improvement. In this paper, we present the discovery that a student model distilled from a few-shot prompted LLM can commonly generalize better than its teacher to unseen examples on such tasks. We find that the student is able to learn a general pattern from the high-quality pseudolabels produced by the teacher during knowledge distillation (KD), and favorably not a general pattern from the low-quality pseudolabels. Leveraging this discovery, we propose a new method, Multistage Collaborative Knowledge Distillation from an LLM (MCKD), for these tasks. MCKD first few-shot prompts an LLM to produce pseudolabels for unlabeled data. Then at each stage of an iterative KD process, a new pair of students is trained on disjoint partitions of the pseudolabeled data, and produces new and improved pseudolabels for their unseen partitions. We conduct extensive experiments on four syntactic and semantic parsing datasets and show the effectiveness of MCKD for lowresource semi-supervised sequence generation. On CRAFT biomedical parsing, for example, 3-stage MCKD with 50 labeled examples outperforms an LLM teacher and vanilla KD by 7.5% and 3.7% parsing F1, respectively, and matches the performance of supervised finetuning with 500 labeled examples. 1 * Equal contribution 1 Our code is made available github.com/andotalao24/Multistage-Collaborative-Knowledge-Distillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Efficient Multivariate Time Series Forecasting via Calibrated Language Models with Privileged Knowledge DistillationChenxi Liu, Hao Miao, Qianxiong Xu, Shaowen Zhou et al.ICDE 2025 · 15 citations
- From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Open-vocabulary Grounded Situation RecognitionChen Cai, Tianyi Liu, Jianjun Gao, Wenyang Liu et al.ACM MM 2025 · 2 citations
- Active Data Curation Effectively Distills Large-Scale Multimodal ModelsVishaal Udandarao, Nikhil Parthasarathy, Muhammad Ferjad Naeem, Talfan Evans et al.CVPR 2025
Builds on9
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled DataColin Wei, Kendrick Shen, Yining Chen, Tengyu MaICLR 2021 · 261 citations
- Large Language Models Are Reasoning TeachersNamgyu Ho, Laura Schmid, Se-Young YunACL 2023 · 102 citations
- Co-training Improves Prompt-based Learning for Large Language ModelsHunter Lang, Monica N. Agrawal, Yoon Kim, David A. SontagICML 2022 · 47 citations
- Does learning require memorization? a short tale about a long tailVitaly FeldmanSTOC 2020 · 28 citations
Related papers
- Structure-Level Knowledge Distillation For Multilingual Sequence LabelingXinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang et al.ACL 2020 · 31 citations
- PromptKD: Unsupervised Prompt Distillation for Vision-Language ModelsZheng Li, Xiang Li, Xinyi Fu, Xin Zhang et al.CVPR 2024
- CLAOCS-TX: Cross-Lingual Triplet Extraction with Aspect-Opinion-Aware Code-Switched Prompting and LLM-Guided Contrastive DistillationLipika Dewangan, Chandresh Kumar MauryaACL 2026
- Multi-level Distillation of Semantic Knowledge for Pre-training Multilingual Language ModelMingqi Li, Fei Ding, Dan Zhang, Long Cheng et al.EMNLP 2022 · 3 citations
- Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved SamplingWenda Xu, Rujun Han, Zifeng Wang, Long T. Le et al.ICLR 2025
