Proper Reuse of Image Classification Features Improves Object Detection
Cristina Nader Vasconcelos, Vighnesh Birodkar, Vincent Dumoulin
摘要
A common practice in transfer learning is to initialize the downstream model weights by pre-training on a data-abundant upstream task. In object detection specifically, the feature backbone is typically initialized with ImageNet classifier weights and fine-tuned on the object detection task. Recent works show this is not strictly necessary under longer training regimes and provide recipes for training the backbone from scratch. We investigate the opposite direction of this end-to-end training trend: we show that an extreme form of knowledge preservation-freezing the classifier-initialized backbone— consistently improves many different detection models, and leads to considerable resource savings. We hypothesize and corroborate experimentally that the remaining detector components capacity and structure is a crucial factor in leveraging the frozen backbone. Immediate applications of our findings include performance improvements on hard cases like detection of long-tail object classes and computational and memory resource savings that contribute to making the field more accessible to researchers with access to fewer computational resources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- SFC: Shared Feature Calibration in Weakly Supervised Semantic SegmentationXinqiao Zhao, Feilong Tang, Xiaoyang Wang, Jimin XiaoAAAI 2024 · 被引用 66 次
- Open-Vocabulary Object Detection upon Frozen Vision and Language ModelsWeicheng Kuo, Yin Cui, Xiuye Gu, A. J. Piergiovanni 等ICLR 2023 · 被引用 37 次
- Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation ModelsShenghao Fu, Junkai Yan, Qize Yang, Xihan Wei 等NeurIPS 2024 · 被引用 24 次
- FedLoGe: Joint Local and Generic Federated Learning under Long-tailed DataZikai Xiao, Zihan Chen, Liyinglan Liu, Yang Feng 等ICLR 2024 · 被引用 14 次
- Could Giant Pre-trained Image Models Extract Universal Representations?Yutong Lin, Ze Liu, Zheng Zhang, Han Hu 等NeurIPS 2022 · 被引用 10 次
它引用的顶会 Paper14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
- Rethinking Pre-training and Self-trainingBarret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui 等NeurIPS 2020 · 被引用 755 次
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 被引用 700 次
相关 Paper
- Distilling Image Classifiers in Object DetectorsShuxuan Guo, José M. Álvarez, Mathieu SalzmannNeurIPS 2021 · 被引用 10 次
- Integrally Migrating Pre-trained Transformer Encoder-decoders for Visual Object DetectionFeng Liu, Xiaosong Zhang, Zhiliang Peng, Zonghao Guo 等ICCV 2023 · 被引用 30 次
- What Makes Instance Discrimination Good for Transfer Learning?Nanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, Stephen LinICLR 2021 · 被引用 183 次
- Training Object Detectors from Scratch: An Empirical Study in the Era of Vision TransformerWeixiang Hong, Jiangwei Lao, Wang Ren, Jian Wang 等CVPR 2022 · 被引用 14 次
- GiraffeDet: A Heavy-Neck Paradigm for Object DetectionYiqi Jiang, Zhiyu Tan, Junyan Wang, Xiuyu Sun 等ICLR 2022 · 被引用 147 次
