Over-parameterized Student Model via Tensor Decomposition Boosted Knowledge Distillation
Yu-Liang Zhan, Zhong-Yi Lu, Hao Sun, Ze-Feng Gao
摘要
Increased training parameters have enabled large pre-trained models to excel in various downstream tasks. Nevertheless, the extensive computational requirements associated with these models hinder their widespread adoption within the community. We focus on Knowledge Distillation (KD), where a compact student model is trained to mimic a larger teacher model, facilitating the transfer of knowledge of large models. In contrast to much of the previous work, we scale up the parameters of the student model during training, to benefit from overparameterization without increasing the inference latency. In particular, we propose a tensor decomposition strategy that effectively over-parameterizes the relatively small student model through an efficient and nearly lossless decomposition of its parameter matrices into higher-dimensional tensors. To ensure efficiency, we further introduce a tensor constraint loss to align the high-dimensional tensors between the student and teacher models. Comprehensive experiments validate the significant performance enhancement by our approach in various KD tasks, covering computer vision and natural language processing areas. Our code is available at https://github.com/intell-sci-comput/OPDF.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Unleashing the Potential of Large Language Models as Prompt Optimizers: Analogical Analysis with Gradient-based Model OptimizersXinyu Tang, Xiaolei Wang, Wayne Xin Zhao, Siyuan Lu 等AAAI 2025 · 被引用 36 次
- Reinforced Informativeness Optimization for Long-Form Retrieval-Augmented GenerationYuhao Wang, Ruiyang Ren, Yucheng Wang, Xin Zhao 等ACL 2026 · 被引用 4 次
- BEE-RAG: Balanced Entropy Engineering for Retrieval-Augmented GenerationYuhao Wang, Ruiyang Ren, Yucheng Wang, Jing Liu 等AAAI 2026 · 被引用 2 次
- L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent InterventionYu-Liang Zhan, Xinyu Tang, Han Wan, Jian Li 等AAAI 2026 · 被引用 2 次
- Unlocking General Long Chain-of-Thought Reasoning Capabilities of Large Language Models via Representation EngineeringXinyu Tang, Xiaolei Wang, Zhihao Lv, Yingqian Min 等ACL 2025
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
相关 Paper
- Knowledge Distillation as Efficient Pre-training: Faster Convergence, Higher Data-efficiency, and Better TransferabilityRuifei He, Shuyang Sun, Jihan Yang, Song Bai 等CVPR 2022 · 被引用 40 次
- Small Pre-trained Language Models Can be Fine-tuned as Large Models via Over-ParameterizationZe-Feng Gao, Kun Zhou, Peiyu Liu, Wayne Xin Zhao 等ACL 2023 · 被引用 3 次
- Exploring extreme parameter compression for pre-trained language modelsBenyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang 等ICLR 2022 · 被引用 23 次
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li 等ACL 2025
- An Empirical Study of Knowledge Distillation for Code Understanding TasksRuiqi Wang, Zezhou Yang, Cuiyun Gao, Xin Xia 等ICSE 2026 · 被引用 1 次
