Over-parameterized Student Model via Tensor Decomposition Boosted Knowledge Distillation
Yu-Liang Zhan, Zhong-Yi Lu, Hao Sun, Ze-Feng Gao
Abstract
Increased training parameters have enabled large pre-trained models to excel in various downstream tasks. Nevertheless, the extensive computational requirements associated with these models hinder their widespread adoption within the community. We focus on Knowledge Distillation (KD), where a compact student model is trained to mimic a larger teacher model, facilitating the transfer of knowledge of large models. In contrast to much of the previous work, we scale up the parameters of the student model during training, to benefit from overparameterization without increasing the inference latency. In particular, we propose a tensor decomposition strategy that effectively over-parameterizes the relatively small student model through an efficient and nearly lossless decomposition of its parameter matrices into higher-dimensional tensors. To ensure efficiency, we further introduce a tensor constraint loss to align the high-dimensional tensors between the student and teacher models. Comprehensive experiments validate the significant performance enhancement by our approach in various KD tasks, covering computer vision and natural language processing areas. Our code is available at https://github.com/intell-sci-comput/OPDF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30c75670-4fb1-4a42-93e6-01f2375e7fd2Cited by top-tier papers5
- Unleashing the Potential of Large Language Models as Prompt Optimizers: Analogical Analysis with Gradient-based Model OptimizersXinyu Tang, Xiaolei Wang, Wayne Xin Zhao, Siyuan Lu et al.AAAI 2025 · 36 citations
- Reinforced Informativeness Optimization for Long-Form Retrieval-Augmented GenerationYuhao Wang, Ruiyang Ren, Yucheng Wang, Xin Zhao et al.ACL 2026 · 4 citations
- BEE-RAG: Balanced Entropy Engineering for Retrieval-Augmented GenerationYuhao Wang, Ruiyang Ren, Yucheng Wang, Jing Liu et al.AAAI 2026 · 2 citations
- L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent InterventionYu-Liang Zhan, Xinyu Tang, Han Wan, Jian Li et al.AAAI 2026 · 2 citations
- Unlocking General Long Chain-of-Thought Reasoning Capabilities of Large Language Models via Representation EngineeringXinyu Tang, Xiaolei Wang, Zhihao Lv, Yingqian Min et al.ACL 2025
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
Related papers
- Knowledge Distillation as Efficient Pre-training: Faster Convergence, Higher Data-efficiency, and Better TransferabilityRuifei He, Shuyang Sun, Jihan Yang, Song Bai et al.CVPR 2022 · 40 citations
- Small Pre-trained Language Models Can be Fine-tuned as Large Models via Over-ParameterizationZe-Feng Gao, Kun Zhou, Peiyu Liu, Wayne Xin Zhao et al.ACL 2023 · 3 citations
- Exploring extreme parameter compression for pre-trained language modelsBenyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang et al.ICLR 2022 · 23 citations
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li et al.ACL 2025
- An Empirical Study of Knowledge Distillation for Code Understanding TasksRuiqi Wang, Zezhou Yang, Cuiyun Gao, Xin Xia et al.ICSE 2026 · 1 citation
