Improving Language Model Distillation through Hidden State Matching
Sayantan Dasgupta, Trevor Cohn
Abstract
Hidden State Matching is shown to improve knowledge distillation of language models by encouraging similarity between a student and its teacher's hidden states, as demonstrated by DistilBERT and its successors. This typically uses a cosine loss, which restricts the dimensionality of the student to the teacher's, severely limiting the compression ratio. We present an alternative technique using Centered Kernel Alignment (CKA) to match hidden states of different dimensionality, allowing for smaller students and higher compression ratios. We show the efficacy of our method using encoder-decoder (BART, mBART & T5) and encoder-only (BERT) architectures across a range of tasks from classification to summarization and translation. Our technique is competitive with the current state-of-the-art distillation methods at comparable compression rates and does not require already pretrained student models. It can scale to students smaller than the current methods, is no slower in training and inference, and is considerably more flexible. The Code is available on github 1 * Also at Google Research 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 593bb5b5-8030-4cd7-bf4f-d46a59223919Cited by top-tier papers6
- Block Recurrent Dynamics in Vision TransformersMozes Jacobs, Thomas Fel, Richard Hakim, Alessandra Brondetta et al.ICLR 2026 · 17 citations
- Don't Ignore the Tail: Decoupled Distillation Produces Top Maths Students on an Academic BudgetSayantan Dasgupta, Trevor Cohn, Tim BaldwinICML 2026
- EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsMinh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep et al.EMNLP 2025
- TALAS: Teacher-Anchored Layer Alignment with Adaptive Sharpness-Aware Minimization for Embedding DistillationQuoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Linh Ngo Van et al.ACL 2026
- WAVE: Window-Aware Vocabulary-Efficient Early-Exit for Training-Free LLM AccelerationSeonggeun Kim, Gilha lee, Hyun KimICML 2026
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 736 citations
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and DepthThao Nguyen, Maithra Raghu, Simon KornblithICLR 2021 · 323 citations
Related papers
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li et al.ACL 2025
- Contrastive Distillation on Intermediate Representations for Language Model CompressionSiqi Sun, Zhe Gan, Yuwei Fang, Yu Cheng et al.EMNLP 2020 · 59 citations
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 8 citations
- Marginal Utility Diminishes: Exploring the Minimum Knowledge for BERT Knowledge DistillationYuanxin Liu, Fandong Meng, Zheng Lin, Weiping Wang et al.ACL 2021
- Knowledge Distillation with the Reused Teacher ClassifierDefang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang et al.CVPR 2022 · 213 citations
