Contrastive Representation Distillation
Yonglong Tian, Dilip Krishnan, Phillip Isola
摘要
Often we wish to transfer representational knowledge from one neural network to another. Examples include distilling a large network into a smaller one, transferring knowledge from one sensory modality to a second, or ensembling a collection of models into a single estimator. Knowledge distillation, the standard approach to these problems, minimizes the KL divergence between the probabilistic outputs of a teacher and student network. We demonstrate that this objective ignores important structural knowledge of the teacher network. This motivates an alternative objective by which we train a student to capture significantly more information in the teacher's representation of the data. We formulate this objective as contrastive learning. Experiments demonstrate that our resulting new objective outperforms knowledge distillation and other cutting-edge distillers on a variety of knowledge transfer tasks, including single model compression, ensemble distillation, and cross-modal transfer. Our method sets a new state-of-the-art in many transfer tasks, and sometimes even outperforms the teacher network when combined with knowledge distillation. Code: this http URL
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper297
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan 等NeurIPS 2020 · 被引用 1,631 次
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 被引用 1,615 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- Long-tailed Recognition by Routing Diverse Distribution-Aware ExpertsXudong Wang, Long Lian, Zhongqi Miao, Ziwei Liu 等ICLR 2021 · 被引用 481 次
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian 等NeurIPS 2022 · 被引用 477 次
它引用的顶会 Paper2
相关 Paper
- Prototypical Contrastive Predictive CodingKyungmin LeeICLR 2022 · 被引用 9 次
- Wasserstein Contrastive Representation DistillationLiqun Chen, Dong Wang, Zhe Gan, Jingjing Liu 等CVPR 2021
- Complementary Relation Contrastive DistillationJinguo Zhu, Shixiang Tang, Dapeng Chen, Shijie Yu 等CVPR 2021
- Contrastive Distillation on Intermediate Representations for Language Model CompressionSiqi Sun, Zhe Gan, Yuwei Fang, Yu Cheng 等EMNLP 2020 · 被引用 59 次
- Knowledge distillation via softmax regression representation learningJing Yang, Brais Martínez, Adrian Bulat, Georgios TzimiropoulosICLR 2021 · 被引用 55 次
