Proper Reuse of Image Classification Features Improves Object Detection
Cristina Nader Vasconcelos, Vighnesh Birodkar, Vincent Dumoulin
Abstract
A common practice in transfer learning is to initialize the downstream model weights by pre-training on a data-abundant upstream task. In object detection specifically, the feature backbone is typically initialized with ImageNet classifier weights and fine-tuned on the object detection task. Recent works show this is not strictly necessary under longer training regimes and provide recipes for training the backbone from scratch. We investigate the opposite direction of this end-to-end training trend: we show that an extreme form of knowledge preservation-freezing the classifier-initialized backbone— consistently improves many different detection models, and leads to considerable resource savings. We hypothesize and corroborate experimentally that the remaining detector components capacity and structure is a crucial factor in leveraging the frozen backbone. Immediate applications of our findings include performance improvements on hard cases like detection of long-tail object classes and computational and memory resource savings that contribute to making the field more accessible to researchers with access to fewer computational resources.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39200a58-60b3-4e92-bb48-2faca881a0ddCited by top-tier papers8
- SFC: Shared Feature Calibration in Weakly Supervised Semantic SegmentationXinqiao Zhao, Feilong Tang, Xiaoyang Wang, Jimin XiaoAAAI 2024 · 66 citations
- Open-Vocabulary Object Detection upon Frozen Vision and Language ModelsWeicheng Kuo, Yin Cui, Xiuye Gu, A. J. Piergiovanni et al.ICLR 2023 · 37 citations
- Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation ModelsShenghao Fu, Junkai Yan, Qize Yang, Xihan Wei et al.NeurIPS 2024 · 24 citations
- FedLoGe: Joint Local and Generic Federated Learning under Long-tailed DataZikai Xiao, Zihan Chen, Liyinglan Liu, Yang Feng et al.ICLR 2024 · 14 citations
- Could Giant Pre-trained Image Models Extract Universal Representations?Yutong Lin, Ze Liu, Zheng Zhang, Han Hu et al.NeurIPS 2022 · 10 citations
Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Rethinking Pre-training and Self-trainingBarret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui et al.NeurIPS 2020 · 755 citations
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 700 citations
Related papers
- Distilling Image Classifiers in Object DetectorsShuxuan Guo, José M. Álvarez, Mathieu SalzmannNeurIPS 2021 · 10 citations
- Integrally Migrating Pre-trained Transformer Encoder-decoders for Visual Object DetectionFeng Liu, Xiaosong Zhang, Zhiliang Peng, Zonghao Guo et al.ICCV 2023 · 30 citations
- What Makes Instance Discrimination Good for Transfer Learning?Nanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, Stephen LinICLR 2021 · 183 citations
- Training Object Detectors from Scratch: An Empirical Study in the Era of Vision TransformerWeixiang Hong, Jiangwei Lao, Wang Ren, Jian Wang et al.CVPR 2022 · 14 citations
- GiraffeDet: A Heavy-Neck Paradigm for Object DetectionYiqi Jiang, Zhiyu Tan, Junyan Wang, Xiuyu Sun et al.ICLR 2022 · 147 citations
