Multi-Mode Online Knowledge Distillation for Self-Supervised Visual Representation Learning
Kaiyou Song, Jin Xie, Shan Zhang, Zimeng Luo
Abstract
Self-supervised learning (SSL) has made remarkable progress in visual representation learning. Some studies combine SSL with knowledge distillation (SSL-KD) to boost the representation learning performance of small models. In this study, we propose a Multi-mode Online Knowledge Distillation method (MOKD) to boost self-supervised visual representation learning. Different from existing SSL-KD methods that transfer knowledge from a static pre-trained teacher to a student, in MOKD, two different models learn collaboratively in a self-supervised manner. Specifically, MOKD consists of two distillation modes: self-distillation and crossdistillation modes. Among them, self-distillation performs self-supervised learning for each model independently, while cross-distillation realizes knowledge interaction between different models. In cross-distillation, a cross-attention feature search strategy is proposed to enhance the semantic feature alignment between different models. As a result, the two models can absorb knowledge from each other to boost their representation learning performance. Extensive experimental results on different backbones and datasets demonstrate that two heterogeneous models can benefit from MOKD and outperform their independently trained baseline. In addition, MOKD also outperforms existing SSL-KD methods for both the student and teacher models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74751e68-a45a-4ab1-b1b9-c24d8e6883a1Cited by top-tier papers6
- Breaking Modality Gap in RGBT Tracking: Coupled Knowledge DistillationAndong Lu, Jiacong Zhao, Chenglong Li, Yun Xiao et al.ACM MM 2024 · 15 citations
- Representing Part-Whole Hierarchies in Foundation Models by Learning Localizability, Composability, and Decomposability from Anatomy via Self-SupervisionMohammad Reza Hosseinzadeh Taher, Michael B. Gotway, Jianming LiangCVPR 2024 · 12 citations
- Learning Task-Agnostic Representations through Multi-Teacher DistillationPhilippe Formont, Maxime Darrin, Banafsheh Karimian, Eric Granger et al.NeurIPS 2025 · 6 citations
- Progressive Mask Distillation for Self-supervised Video RepresentationKewei Wu, Chong Liang, Zhao Xie, Dan GuoCVPR 2026
- Foundation-Adaptive Integrated Refinement for Generalized Category DiscoveryYuwei Bian, Shidong Wang, Yazhou Yao, Haofeng ZhangAAAI 2026
Builds on29
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Unsupervised Representation Transfer for Small Networks: I Believe I Can Distill On-the-FlyHee Min Choi, Hyoa Kang, Dokwan OhNeurIPS 2021 · 13 citations
- C2KD: Bridging the Modality Gap for Cross-Modal Knowledge DistillationFushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang et al.CVPR 2024
- ScaleKD: Distilling Scale-Aware Knowledge in Small Object DetectorYichen Zhu, Qiqi Zhou, Ning Liu, Zhiyuan Xu et al.CVPR 2023
- MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge DistillationHui Li, Pengfei Yang, Juanyang Chen, Le Dong et al.ACM MM 2025 · 4 citations
- Refine Myself by Teaching Myself: Feature Refinement via Self-Knowledge DistillationMingi Ji, Seungjae Shin, Seunghyun Hwang, Gibeom Park et al.CVPR 2021
