Learning Student Networks with Few Data
Shumin Kong, Tianyu Guo, Shan You, Chang Xu
摘要
Recently, the teacher-student learning paradigm has drawn much attention in compressing neural networks on low-end edge devices, such as mobile phones and wearable watches. Current algorithms mainly assume the complete dataset for the teacher network is also available for the training of the student network. However, for real-world scenarios, users may only have access to part of training examples due to commercial profits or data privacy, and severe over-fitting issues would happen as a result. In this paper, we tackle the challenge of learning student networks with few data by investigating the ground-truth data-generating distribution underlying these few data. Taking Wasserstein distance as the measurement, we assume this ideal data distribution lies in a neighborhood of the discrete empirical distribution induced by the training examples. Thus we propose to safely optimize the worst-case cost within this neighborhood to boost the generalization. Furthermore, with theoretical analysis, we derive a novel and easy-to-implement loss for training the student network in an end-to-end fashion. Experimental results on benchmark datasets validate the effectiveness of our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian 等NeurIPS 2022 · 被引用 477 次
- SCOP: Scientific Control for Reliable Neural Network PruningYehui Tang, Yunhe Wang, Yixing Xu, Dacheng Tao 等NeurIPS 2020 · 被引用 208 次
- Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient SpaceShangchen Du, Shan You, Xiaojie Li, Jianlong Wu 等NeurIPS 2020 · 被引用 144 次
- Locally Free Weight Sharing for Network Width SearchXiu Su, Shan You, Tao Huang, Fei Wang 等ICLR 2021 · 被引用 45 次
- GreedyNAS: Towards Fast One-Shot NAS With Greedy SupernetShan You, Tao Huang, Mingmin Yang, Fei Wang 等CVPR 2020
相关 Paper
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 被引用 66 次
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang 等ICCV 2019 · 被引用 427 次
- Data-Free Ensemble Knowledge Distillation for Privacy-conscious Multimedia Model CompressionZhiwei Hao, Yong Luo, Han Hu, Jianping An 等ACM MM 2021 · 被引用 11 次
- Distilling Portable Generative Adversarial Networks for Image TranslationHanting Chen, Yunhe Wang, Han Shu, Changyuan Wen 等AAAI 2020 · 被引用 89 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
