Channel-wise Knowledge Distillation for Dense Prediction*
Changyong Shu, Yifan Liu, Jianfei Gao, Zheng Yan, Chunhua Shen
Abstract
Knowledge distillation (KD) has been proven a simple and effective tool for training compact dense prediction models. Lightweight student networks are trained by extra supervision transferred from large teacher networks. Most previous KD variants for dense prediction tasks align the activation maps from the student and teacher network in the spatial domain, typically by normalizing the activation values on each spatial location and minimizing point-wise and/or pair-wise discrepancy. Different from the previous methods, here we propose to normalize the activation map of each channel to obtain a soft probability map. By simply minimizing the Kullback–Leibler (KL) divergence between the channel-wise probability map of the two networks, the distillation process pays more attention to the most salient regions of each channel, which are valuable for dense prediction tasks.We conduct experiments on a few dense prediction tasks, including semantic segmentation and object detection. Experiments demonstrate that our proposed method outperforms state-of-the-art distillation methods considerably, and can require less computational cost during training. In particular, we improve the RetinaNet detector (ResNet50 backbone) by 3.4% in mAP on the COCO dataset, and PSPNet (ResNet18 backbone) by 5.81% in mIoU on the Cityscapes dataset. Code is available at: https://git.io/Distiller
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9fe8c625-b888-47ef-a2fc-92e0256566b6Cited by top-tier papers61
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian et al.NeurIPS 2022 · 477 citations
- Cross-Image Relational Knowledge Distillation for Semantic SegmentationChuanguang Yang, Helong Zhou, Zhulin An, Xue Jiang et al.CVPR 2022 · 228 citations
- Knowledge Distillation with the Reused Teacher ClassifierDefang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang et al.CVPR 2022 · 213 citations
- SCTNet: Single-Branch CNN with Transformer Semantic Information for Real-Time SegmentationZhengze Xu, Dongyue Wu, Changqian Yu, Xiangxiang Chu et al.AAAI 2024 · 166 citations
Builds on10
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- RepPoints: Point Set Representation for Object DetectionZe Yang, Shaohui Liu, Han Hu, Liwei Wang et al.ICCV 2019 · 1,056 citations
- Asymmetric Non-Local Neural Networks for Semantic SegmentationZhen Zhu, Mengdu Xu, Song Bai, Tengteng Huang et al.ICCV 2019 · 694 citations
- Deep Multimodal Fusion by Channel ExchangingYikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu et al.NeurIPS 2020 · 321 citations
- Improve Object Detection with Feature-based Knowledge Distillation: Towards Accurate and Efficient DetectorsLinfeng Zhang, Kaisheng MaICLR 2021 · 251 citations
Related papers
- DCSF-KD: Dynamic Channel-wise Spatial Feature Knowledge Distillation for Object DetectionTao Dai, Yang Lin, Hang Guo, Jinbao Wang et al.AAAI 2025 · 7 citations
- Knowledge Distillation for Object Detection via Rank Mimicking and Prediction-Guided Feature ImitationGang Li, Xiang Li, Yujie Wang, Shanshan Zhang et al.AAAI 2022 · 105 citations
- Pixel-Wise Contrastive DistillationJunqiang Huang, Zichao GuoICCV 2023 · 8 citations
- Bridging Cross-task Protocol Inconsistency for Distillation in Dense Object DetectionLongrong Yang, Xianpan Zhou, Xuewei Li, Liang Qiao et al.ICCV 2023 · 51 citations
- Localization Distillation for Dense Object DetectionZhaohui Zheng, Rongguang Ye, Ping Wang, Dongwei Ren et al.CVPR 2022 · 177 citations
