A Good Student is Cooperative and Reliable: CNN-Transformer Collaborative Learning for Semantic Segmentation
Jinjing Zhu, Yunhao Luo, Xu Zheng, Hao Wang, Lin Wang
Abstract
In this paper, we strive to answer the question ‘how to collaboratively learn convolutional neural network (CNN)-based and vision transformer (ViT)-based models by selecting and exchanging the reliable knowledge between them for semantic segmentation?’ Accordingly, we propose an online knowledge distillation (KD) framework that can simultaneously learn compact yet effective CNN-based and ViT-based models with two key technical breakthroughs to take full advantage of CNNs and ViT while compensating their limitations. Firstly, we propose heterogeneous feature distillation (HFD) to improve students’ consistency in low-layer feature space by mimicking heterogeneous features between CNNs and ViT. Secondly, to facilitate the two students to learn reliable knowledge from each other, we propose bidirectional selective distillation (BSD) that can dynamically transfer selective knowledge. This is achieved by 1) region-wise BSD determining the directions of knowledge transferred between the corresponding regions in the feature space and 2) pixel-wise BSD discerning which of the prediction knowledge to be transferred in the logit space. Extensive experiments on three benchmark datasets demonstrate that our proposed framework outperforms the state-of-the-art online distillation methods by a large margin, and shows its efficacy in learning collaboratively between ViT-based and CNN-based models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- SCTNet: Single-Branch CNN with Transformer Semantic Information for Real-Time SegmentationZhengze Xu, Dongyue Wu, Changqian Yu, Xiangxiang Chu et al.AAAI 2024 · 166 citations
- ScaleKD: Strong Vision Transformers Could Be Excellent TeachersJiawei Fan, Chao Li, Xiaolong Liu, Anbang YaoNeurIPS 2024 · 20 citations
- Semantics, Distortion, and Style Matter: Towards Source-Free UDA for Panoramic SegmentationXu Zheng, Pengyuan Zhou, Athanasios V. Vasilakos, Lin WangCVPR 2024 · 16 citations
- GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-Aware Panoramic Semantic SegmentationWeiming Zhang, Yexin Liu, Xu Zheng, Lin WangCVPR 2024 · 14 citations
- Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse DistillationJinjing Zhu, Tianbo Pan, Zidong Cao, Yexin Liu et al.ICCV 2025 · 3 citations
Builds on16
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped WindowsXiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang et al.CVPR 2022 · 1,207 citations
Related papers
- Distilling Knowledge from Heterogeneous Architectures for Semantic SegmentationYanglin Huang, Kai Hu, Yuan Zhang, Zhineng Chen et al.AAAI 2025 · 4 citations
- Heterogeneous Complementary DistillationLiuchi Xu, Hao Zheng, Lu Wang, Lisheng Xu et al.AAAI 2026
- Fuse Before Transfer: Knowledge Fusion for Heterogeneous DistillationGuopeng Li, Qiang Wang, Ke Yan, Shouhong Ding et al.ICCV 2025 · 1 citation
- Cumulative Spatial Knowledge Distillation for Vision TransformersBorui Zhao, Renjie Song, Jiajun LiangICCV 2023 · 27 citations
- Generic-to-Specific Distillation of Masked AutoencodersWei Huang, Zhiliang Peng, Li Dong, Furu Wei et al.CVPR 2023
