Lion: Adversarial Distillation of Proprietary Large Language Models
Yuxin Jiang, Chunkit Chan, Mingyang Chen, Wei Wang
摘要
The practice of transferring knowledge from a sophisticated, proprietary large language model (LLM) to a compact, open-source LLM has garnered considerable attention. Previous works have focused on a unidirectional knowledge distillation way by aligning the responses of the student model with those of the teacher models to a set of instructions. Nevertheless, they overlooked the possibility of incorporating any "feedback"-identifying challenging instructions where the student model's performance falls short-to boost the student model's proficiency iteratively. To this end, we propose a novel adversarial distillation framework for a more efficient knowledge transfer. Leveraging the versatile role adaptability of LLMs, we prompt the teacher model to identify "hard" instructions and generate new "hard" instructions for the student model, creating a three-stage adversarial loop of imitation, discrimination, and generation. By applying this adversarial framework, we successfully transfer knowledge from ChatGPT to a student model (named Lion), using a mere 70k training data. Our results show that Lion-13B not only achieves comparable open-ended generation capabilities to Chat-GPT but surpasses conventional state-of-the-art (SOTA) instruction-tuned models like Vicuna-13B by 55.4% in challenging zero-shot reasoning benchmarks such as BIG-Bench Hard (BBH) and 16.7% on AGIEval. 1 * The two authors have equal contributions. 1 Code and model can be found at https://github.com/ YJiangcm/Lion .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang 等NeurIPS 2024 · 被引用 50 次
- FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language ModelsYuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong 等ACL 2024 · 被引用 10 次
- Learning to Edit: Aligning LLMs with Knowledge EditingYuxin Jiang, Yufei Wang, Chuhan Wu, Wanjun Zhong 等ACL 2024 · 被引用 9 次
- xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation MetricsDaniil Larionov, Mikhail Seleznyov, Vasiliy Viskov, Alexander Panchenko 等EMNLP 2024 · 被引用 2 次
- ECON: On the Detection and Resolution of Evidence ConflictsCheng Jiayang, Chunkit Chan, Qianqian Zhuang, Lin Qiu 等EMNLP 2024 · 被引用 2 次
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 被引用 887 次
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao 等ICLR 2022 · 被引用 237 次
相关 Paper
- Personalized Distillation: Empowering Open-Sourced LLMs with Adaptive Learning for Code GenerationHailin Chen, Amrita Saha, Steven Chu-Hong Hoi, Shafiq JotyEMNLP 2023 · 被引用 8 次
- Pedagogically-Inspired Data Synthesis for Language Model Knowledge DistillationBowei He, Yankai Chen, Xiaokun Zhang, Linghe Kong 等ICLR 2026 · 被引用 2 次
- Probing to Refine: Reinforcement Distillation of LLM Reasoners via Explanatory InversionZhen Tan, Chengshuai Zhao, Song Wang, Jundong Li 等ICLR 2026
- QCRD: Quality-guided Contrastive Rationale Distillation for Large Language ModelsWei Wang, Zhaowei Li, Qi Xu, Yiqing Cai 等EMNLP 2025 · 被引用 1 次
- Towards Zero-Shot Knowledge Distillation for Natural Language ProcessingAhmad Rashid, Vasileios Lioutas, Abbas Ghaddar, Mehdi RezagholizadehEMNLP 2021 · 被引用 25 次
