Adversarial Moment-Matching Distillation of Large Language Models
Chen Jia
摘要
Knowledge distillation (KD) has been shown to be highly effective in guiding a student model with a larger teacher model and achieving practical benefits in improving the computational and memory efficiency for large language models (LLMs). State-of-the-art KD methods for LLMs mostly rely on minimizing explicit distribution distance between teacher and student probability predictions. Instead of optimizing these mandatory behaviour cloning objectives, we explore an imitation learning strategy for KD of LLMs. In particular, we minimize the imitation gap by matching the action-value moments of the teacher's behavior from both on- and off-policy perspectives. To achieve this action-value moment-matching goal, we propose an adversarial training algorithm to jointly estimate the moment-matching distance and optimize the student policy to minimize it. Results from both task-agnostic instruction-following experiments and task-specific experiments demonstrate the effectiveness of our method and achieve new state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- UORA: Uniform Orthogonal Reinitialization Adaptation in Parameter Efficient Fine-Tuning of Large ModelsXueyan Zhang, Jinman Zhao, Zhifei Yang, Yibo Zhong 等ACL 2025 · 被引用 15 次
- Text Generation as Continuous Latent Dynamics via Reinforcement LearningChen JiaICML 2026
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- On-Policy Distillation of Language Models: Learning from Self-Generated MistakesRishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk 等ICLR 2024 · 被引用 311 次
- Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP TasksYizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi 等EMNLP 2022 · 被引用 238 次
- Less is More: Task-aware Layer-wise Distillation for Language Model CompressionChen Liang, Simiao Zuo, Qingru Zhang, Pengcheng He 等ICML 2023 · 被引用 119 次
相关 Paper
- DistiLLM: Towards Streamlined Distillation for Large Language ModelsJongwoo Ko, Sungnyun Kim, Tianyi Chen, Se-Young YunICML 2024 · 被引用 86 次
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 被引用 95 次
- Revisiting Knowledge Distillation for Autoregressive Language ModelsQihuang Zhong, Liang Ding, Li Shen, Juhua Liu 等ACL 2024
- Why Knowledge Distillation Works in Generative Models: A Minimal Working ExplanationSungmin Cha, Kyunghyun ChoNeurIPS 2025 · 被引用 9 次
- Language Model Distillation: A Temporal Difference Imitation Learning PerspectiveZishun Yu, Shangzhe Li, Xinhua ZhangAAAI 2026
