Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object Detection
Yuxin Fang, Shusheng Yang, Shijie Wang, Yixiao Ge, Ying Shan, Xinggang Wang
摘要
We present an approach to efficiently and effectively adapt a masked image modeling (MIM) pre-trained vanilla Vision Transformer (ViT) for object detection, which is based on our two novel observations: (i) A MIM pre-trained vanilla ViT encoder can work surprisingly well in the challenging object-level recognition scenario even with randomly sampled partial observations, e.g., only 25% 50% of the input embeddings. (ii) In order to construct multi-scale representations for object detection from single-scale ViT, a randomly initialized compact convolutional stem supplants the pre-trained patchify stem, and its intermediate features can naturally serve as the higher resolution inputs of a feature pyramid network without further upsampling or other manipulations. While the pre-trained ViT is only regarded as the 3rd-stage of our detector’s backbone instead of the whole feature extractor. This naturally results in a ConvNet-ViT hybrid architecture. The proposed detector, named MimDet, enables a MIM pre-trained vanilla ViT to outperform leading hierarchical architectures such as Swin Transformer, MViTv2 and ConvNeXt on COCO object detection & instance segmentation, and achieves better results compared with the previous best adapted vanilla ViT detector using a more modest fine-tuning recipe while converging 2.8× faster. Code and pre-trained models are available at https://github.com/hustvl/MIMDet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He 等ICLR 2023 · 被引用 204 次
- MCMAE: Masked Convolution Meets Masked AutoencodersPeng Gao, Teli Ma, Hongsheng Li, Ziyi Lin 等NeurIPS 2022 · 被引用 84 次
- Spatial Transform Decoupling for Oriented Object DetectionHongtian Yu, Yunjie Tian, Qixiang Ye, Yunfan LiuAAAI 2024 · 被引用 56 次
- Integrally Migrating Pre-trained Transformer Encoder-decoders for Visual Object DetectionFeng Liu, Xiaosong Zhang, Zhiliang Peng, Zonghao Guo 等ICCV 2023 · 被引用 30 次
- DETR Does Not Need Multi-Scale or Locality DesignYutong Lin, Yuhui Yuan, Zheng Zhang, Chen Li 等ICCV 2023 · 被引用 27 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- HiViT: A Simpler and More Efficient Design of Hierarchical Vision TransformerXiaosong Zhang, Yunjie Tian, Lingxi Xie, Wei Huang 等ICLR 2023
- Global Context Vision TransformersAli Hatamizadeh, Hongxu Yin, Greg Heinrich, Jan Kautz 等ICML 2023 · 被引用 213 次
- Masked Image Residual Learning for Scaling Deeper Vision TransformersGuoxi Huang, Hongtao Fu, Adrian G. BorsNeurIPS 2023 · 被引用 10 次
- Masked Image Modeling with Denoising ContrastKun Yi, Yixiao Ge, Xiaotong Li, Shusheng Yang 等ICLR 2023 · 被引用 8 次
- Integrally Pre-Trained Transformer Pyramid NetworksYunjie Tian, Lingxi Xie, Zhaozhi Wang, Longhui Wei 等CVPR 2023
