Knowledge Distillation via the Target-aware Transformer
Sihao Lin, Hongwei Xie, Bing Wang, Kaicheng Yu, Xiaojun Chang, Xiaodan Liang, Gang Wang
摘要
Knowledge distillation becomes a de facto standard to improve the performance of small neural networks. Most of the previous works propose to regress the representational features from the teacher to the student in a oneto-one spatial matching fashion. However, people tend to overlook the fact that, due to the architecture differences, the semantic information on the same spatial location usually vary. This greatly undermines the underlying assumption of the one-to-one distillation approach. To this end, we propose a novel one-to-all spatial matching knowledge distillation approach. Specifically, we allow each pixel of the teacher feature to be distilled to all spatial locations of the student features given its similarity, which is generated from a target-aware transformer. Our approach surpasses the state-of-the-art methods by a significant margin on various computer vision benchmarks, such as Im-ageNet, Pascal VOC and COCOStuff10k. Code is available at https://github.com/sihaoevery/TaT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Logit Standardization in Knowledge DistillationShangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang 等CVPR 2024 · 被引用 183 次
- From Knowledge Distillation to Self-Knowledge Distillation: A Unified Approach with Normalized Loss and Customized Soft LabelsZhendong Yang, Ailing Zeng, Zhe Li, Tianke Zhang 等ICCV 2023 · 被引用 141 次
- Cross-modal Clinical Graph Transformer for Ophthalmic Report GenerationMingjie Li, Wenjia Cai, Karin Verspoor, Shirui Pan 等CVPR 2022 · 被引用 55 次
- Mask Propagation for Efficient Video Semantic SegmentationYuetian Weng, Mingfei Han, Haoyu He, Mingjie Li 等NeurIPS 2023 · 被引用 36 次
- DomainAdaptor: A Novel Approach to Test-time AdaptationJian Zhang, Lei Qi, Yinghuan Shi, Yang GaoICCV 2023 · 被引用 31 次
它引用的顶会 Paper16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
相关 Paper
- Distilling Knowledge from Heterogeneous Architectures for Semantic SegmentationYanglin Huang, Kai Hu, Yuan Zhang, Zhineng Chen 等AAAI 2025 · 被引用 4 次
- Distilling Global and Local Logits with Densely Connected RelationsYoumin Kim, Jinbae Park, Younho Jang, Muhammad Salman Ali 等ICCV 2021 · 被引用 33 次
- DETRDistill: A Universal Knowledge Distillation Framework for DETR-familiesJiahao Chang, Shuo Wang, Hai-Ming Xu, Zehui Chen 等ICCV 2023 · 被引用 53 次
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian 等NeurIPS 2022 · 被引用 477 次
- Cross-Image Relational Knowledge Distillation for Semantic SegmentationChuanguang Yang, Helong Zhou, Zhulin An, Xue Jiang 等CVPR 2022 · 被引用 228 次
