Multi-Label Knowledge Distillation
Penghui Yang, Ming-Kun Xie, Chen-Chen Zong, Lei Feng, Gang Niu, Masashi Sugiyama, Sheng-Jun Huang
Abstract
Existing knowledge distillation methods typically work by imparting the knowledge of output logits or intermediate feature maps from the teacher network to the student network, which is very successful in multi-class single-label learning. However, these methods can hardly be extended to the multi-label learning scenario, where each instance is associated with multiple semantic labels, because the prediction probabilities do not sum to one and feature maps of the whole example may ignore minor classes in such a scenario. In this paper, we propose a novel multi-label knowledge distillation method. On one hand, it exploits the informative semantic knowledge from the logits by dividing the multi-label learning problem into a set of binary classification problems; on the other hand, it enhances the distinctiveness of the learned feature representations by leveraging the structural information of label-wise embeddings. Experimental results on multiple benchmark datasets validate that the proposed method can avoid knowledge counteraction among labels, thus achieving superior performance against diverse comparing methods. Our code is available at: https://github.com/penghui-yang/L2D .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7675b86c-8713-4bf5-8205-d763eb0723d2Cited by top-tier papers9
- Robust Synthetic-to-Real Transfer for Stereo MatchingJiawei Zhang, Jiahe Li, Lei Huang, Xiaohan Yu et al.CVPR 2024 · 12 citations
- Multi-View Multi-Label Classification via View-Label Matching SelectionHao Wei, Yongjian Deng, Qiuru Hai, Yuena Lin et al.AAAI 2025 · 3 citations
- FedHarmony: Harmonizing Heterogeneous Label Correlations in Federated Multi-Label LearningZhiqiang Kou, Junxiang Wu, Wenke Huang, Wenwen He et al.CVPR 2026 · 3 citations
- Multi-label Self Knowledge DistillationXucong Wang, Pengkun Wang, Shurui Zhang, Miao Fang et al.AAAI 2025 · 2 citations
- Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary HeadPenghui Yang, Chen-Chen Zong, Sheng-Jun Huang, Lei Feng et al.KDD 2025 · 1 citation
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
Related papers
- Multi-Level Logit DistillationYing Jin, Jiaqi Wang, Dahua LinCVPR 2023
- Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge DistillationJiaming Lv, Haoyuan Yang, Peihua LiNeurIPS 2024 · 59 citations
- Scale Decoupled DistillationShicai Wei, Chunbo Luo, Yang LuoCVPR 2024 · 32 citations
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian et al.NeurIPS 2022 · 477 citations
- Complementary Relation Contrastive DistillationJinguo Zhu, Shixiang Tang, Dapeng Chen, Shijie Yu et al.CVPR 2021
