Pay Attention to Features, Transfer Learn Faster CNNs
Kafeng Wang, Xitong Gao, Yiren Zhao, Xingjian Li, Dejing Dou, Cheng-Zhong Xu
Abstract
Deep convolutional neural networks are now widely deployed in vision applications, but a limited size of training data can restrict their task performance. Transfer learning offers the chance for CNNs to learn with limited data samples by transferring knowledge from models pretrained on large datasets. Blindly transferring all learned features from the source dataset, however, brings unnecessary computation to CNNs on the target task. In this paper, we propose attentive feature distillation and selection (AFDS), which not only adjusts the strength of transfer learning regularization but also dynamically determines the important features to transfer. By deploying AFDS on ResNet-101, we achieved a state-of-the-art computation reduction at the same accuracy budget, outperforming all existing transfer learning methods. With a 10x MACs reduction budget, a ResNet-101 equipped with AFDS transfer learned from ImageNet to Stanford Dogs 120, can achieve an accuracy 11.07% higher than its best competitor.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers7
- Unraveling Meta-Learning: Understanding Feature Representations for Few-Shot TasksMicah Goldblum, Steven Reich, Liam Fowl, Renkun Ni et al.ICML 2020 · 82 citations
- Task-Oriented Feature DistillationLinfeng Zhang, Yukang Shi, Zuoqiang Shi, Kaisheng Ma et al.NeurIPS 2020 · 74 citations
- Lipschitz Continuity Guided Knowledge DistillationYuzhang Shang, Bin Duan, Ziliang Zong, Liqiang Nie et al.ICCV 2021 · 31 citations
- Multi-Knowledge Aggregation and Transfer for Semantic SegmentationYuang Liu, Wei Zhang, Jun WangAAAI 2022 · 11 citations
- Mitigating Semantic Collapse in Partially Relevant Video RetrievalWonJun Moon, Minseok Jung, Gilhan Park, Tae-Young Kim et al.NeurIPS 2025 · 7 citations
Related papers
- Transferring Knowledge From Large Foundation Models to Small Downstream ModelsShikai Qiu, Boran Han, Danielle C. Maddix, Shuai Zhang et al.ICML 2024 · 9 citations
- Show, Attend and Distill: Knowledge Distillation via Attention-based Feature MatchingMingi Ji, Byeongho Heo, Sungrae ParkAAAI 2021 · 194 citations
- Distilling from Similar Tasks for Transfer Learning on a BudgetKenneth Borup, Cheng Perng Phoo, Bharath HariharanICCV 2023 · 3 citations
- Learning an Inference-accelerated Network from a Pre-trained Model with Frequency-enhanced Feature DistillationXuesong Niu, Jili Gu, Guoxin Zhang, Pengfei Wan et al.ACM MM 2022 · 1 citation
- AdaFilter: Adaptive Filter Fine-Tuning for Deep Transfer LearningYunhui Guo, Yandong Li, Liqiang Wang, Tajana RosingAAAI 2020 · 44 citations
