The Bridge-Garden Dilemma in LLM Distillation: Why Mixing Hard and Soft Labels Works
Guanghui Wang, Kaiwen Kacuila, zhiyong yang, Zitai Wang, Jin-Wen Wu, Longtao Huang, Qianqian Xu, Qingming Huang
Abstract
Knowledge distillation (KD) transfers knowledge from a large teacher model to a smaller student. In language modeling, the student is trained either on tokens sampled from the teacher (hard labels) or the teacher’s full next-token distribution (soft labels). Despite soft labels appear strictly richer, we find that mixing hard and soft labels consistently yields better results. Crucially, we show that this gain cannot be explained by closer teacher matching during training. Instead, it comes from reduced exposure bias---the mismatch between training and inference distributions. To explain this phenomenon, we introduce the Bridge–Garden Decomposition theory, which categorizes generation steps into two types: Bridges, where the next token must be exact, and Gardens, where it can be flexible. We show that hard-only KD excels in Bridges by avoiding risky deviations, while soft-only KD preserves diversity in Gardens. A hybrid strategy handles both cases and, as a result, reduces exposure bias across the sequence. Guided by this theory, we develop a family of Bridge--Garden hybrid supervision methods that adaptively balance hard and soft labels. Across seven teacher--student pairs (including Qwen, Llama, Gemma, and DeepSeek) and benchmarks in reasoning and coding, our approach outperforms divergence-based and on-policy KD baselines while reducing training cost by 9.7, enabling efficient model compression.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93b4d3f3-83dc-45cf-924e-28bdcbe5eca1Builds on29
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu et al.ICLR 2024 · 637 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
Related papers
- Hybrid Policy Distillation for LLMsWenhong Zhu, Ruobing Xie, Rui Wang, Pengfei LiuICML 2026 · 2 citations
- ABKD: Pursuing a Proper Allocation of the Probability Mass in Knowledge Distillation via α-β-DivergenceGuanghui Wang, Zhiyong Yang, Zitai Wang, Shi Wang et al.ICML 2025
- Revisiting Knowledge Distillation for Autoregressive Language ModelsQihuang Zhong, Liang Ding, Li Shen, Juhua Liu et al.ACL 2024
- Can Students Beyond the Teacher? Distilling Knowledge from Teacher's BiasJianhua Zhang, Yi Gao, Ruyu Liu, Xu Cheng et al.AAAI 2025 · 2 citations
- GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale InjectionHyeongsoo Lim, Hyung Yong Kim, Jin Young Kim, Min Ho Jang et al.AAAI 2026
