Wasserstein Contrastive Representation Distillation
Liqun Chen, Dong Wang, Zhe Gan, Jingjing Liu, Ricardo Henao, Lawrence Carin
Abstract
The primary goal of knowledge distillation (KD) is to encapsulate the information of a model learned from a teacher network into a student network, with the latter being more compact than the former. Existing work, e.g., using Kullback-Leibler divergence for distillation, may fail to capture important structural knowledge in the teacher network and often lacks the ability for feature generalization, particularly in situations when teacher and student are built to address different classification tasks. We propose Wasserstein Contrastive Representation Distillation (WCoRD), which leverages both primal and dual forms of Wasserstein distance for KD. The dual form is used for global knowledge transfer, yielding a contrastive learning objective that maximizes the lower bound of mutual information between the teacher and the student networks. The primal form is used for local contrastive knowledge transfer within a mini-batch, effectively matching the distributions of features between the teacher and the student networks. Experiments demonstrate that the proposed WCoRD method outperforms state-of-the-art approaches on privileged information distillation, model compression and cross-modal transfer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers25
- Compressing Visual-linguistic Model via Knowledge DistillationZhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lijuan Wang et al.ICCV 2021 · 121 citations
- Modality-aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence DetectionJiashuo Yu, Jinyu Liu, Ying Cheng, Rui Feng et al.ACM MM 2022 · 64 citations
- Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge DistillationJiaming Lv, Haoyuan Yang, Peihua LiNeurIPS 2024 · 59 citations
- Data Efficient Language-Supervised Zero-Shot Recognition with Optimal Transport DistillationBichen Wu, Ruizhe Cheng, Peizhao Zhang, Tianren Gao et al.ICLR 2022 · 57 citations
- DeepWSD: Projecting Degradations in Perceptual Space to Wasserstein Distance in Deep Feature SpaceXingran Liao, Baoliang Chen, Hanwei Zhu, Shiqi Wang et al.ACM MM 2022 · 32 citations
Builds on11
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
Related papers
- Complementary Relation Contrastive DistillationJinguo Zhu, Shixiang Tang, Dapeng Chen, Shijie Yu et al.CVPR 2021
- Prototypical Contrastive Predictive CodingKyungmin LeeICLR 2022 · 9 citations
- Pay Attention to Your Positive Pairs: Positive Pair Aware Contrastive Knowledge DistillationZhipeng Yu, Qianqian Xu, Yangbangyan Jiang, Haoyu Qin et al.ACM MM 2022 · 10 citations
- Enhanced Multimodal Representation Learning with Cross-modal KDMengxi Chen, Linyu Xing, Yu Wang, Ya ZhangCVPR 2023
- MCW-KD: Multi-Cost Wasserstein Knowledge Distillation for Large Language ModelsHoang Tran Vuong, Tue Le, Quyen Tran, Linh Ngo Van et al.AAAI 2026
