Personalized Distillation: Empowering Open-Sourced LLMs with Adaptive Learning for Code Generation
Hailin Chen, Amrita Saha, Steven Chu-Hong Hoi, Shafiq Joty
Abstract
With the rise of powerful closed-sourced LLMs (ChatGPT, GPT-4), there are increasing interests in distilling the capabilies of close-sourced LLMs to smaller open-sourced LLMs. Previous distillation methods usually prompt Chat-GPT to generate a set of instructions and answers, for the student model to learn. However, such standard distillation approach neglects the merits and conditions of the student model. Inspired by modern teaching principles, we design a personalised distillation process, in which the student attempts to solve a task first, then the teacher provides an adaptive refinement for the student to improve. Instead of feeding the student with teacher's prior, personalised distillation enables personalised learning for the student model, as it only learns on examples it makes mistakes upon and learns to improve its own solution. On code generation, personalised distillation consistently outperforms standard distillation with only one third of the data. With only 2.5-3K personalised examples that incur a data-collection cost of 4-6$, we boost CodeGen-mono-16B by 7% to achieve 36.4% pass@1 and StarCoder by 12.2% to achieve 45.8% pass@1 on HumanEval. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Faster Speculative Decoding via Effective Draft Decoder with Pruned Candidate TreeHuanran Zheng, Xiaoling WangACL 2025 · 6 citations
- Smaller but Better: Self-Paced Knowledge Distillation for Lightweight yet Effective LCMsYujia Chen, Yang Ye, Zhongqi Li, Yuchi Ma et al.FSE 2025 · 1 citation
- Automated Mass Malware Factory: The Convergence of Piggybacking and Adversarial Example in Android Malicious Software GenerationHeng Li, Zhiyuan Yao, Bang Wu, Cuiying Gao et al.NDSS 2025
- Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV OptimizationGuanghui Song, Dongping Liao, Yiren Zhao, Kejiang Ye et al.EMNLP 2025
- AP2O-Coder: Adaptively Progressive Preference Optimization for Reducing Compilation and Runtime Errors in LLM-Generated CodeJianqing Zhang, Wei Xia, Hande Dong, Qiang Lin et al.AAAI 2026
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
Related papers
- AMR-Evol: Adaptive Modular Response Evolution Elicits Better Knowledge Distillation for Large Language Models in Code GenerationZiyang Luo, Xin Li, Hongzhan Lin, Jing Ma et al.EMNLP 2024 · 1 citation
- SelfCodeAlign: Self-Alignment for Code GenerationYuxiang Wei, Federico Cassano, Jiawei Liu, Yifeng Ding et al.NeurIPS 2024 · 79 citations
- Lion: Adversarial Distillation of Proprietary Large Language ModelsYuxin Jiang, Chunkit Chan, Mingyang Chen, Wei WangEMNLP 2023 · 20 citations
- Democratizing Reasoning Ability: Tailored Learning from Large Language ModelZhaoyang Wang, Shaohan Huang, Yuxuan Liu, Jiahai Wang et al.EMNLP 2023 · 8 citations
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen et al.ICLR 2024 · 118 citations
