Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling
Keyu Tian, Yi Jiang, Qishuai Diao, Chen Lin, Liwei Wang, Zehuan Yuan
摘要
We identify and overcome two key obstacles in extending the success of BERT-style pre-training, or masked image modeling, to convolutional networks (convnets): (i) convolution operation cannot handle irregular, randomly masked input images; (ii) the single-scale nature of BERT pre-training is inconsistent with convnet's hierarchical structure. For (i), we treat unmasked pixels as sparse voxels of 3D point clouds and use sparse convolution to encode. This is the first use of sparse convolution for 2D masked modeling. For (ii), we develop a hierarchical decoder to reconstruct images from multi-scale encoded features. Our method, called Sparse masKed modeling (SparK), is general: it can be used directly on any convolutional model without backbone modifications. We validate it on both classical (ResNet) and modern (ConvNeXt) models: on three downstream tasks, it surpasses both state-of-the-art contrastive learning and transformer-based masked modeling by similarly large margins (around +1.0%). The improvements on object detection and instance segmentation are more significant (up to +3.5%), validating the strong transferability of features learned. We also find its favorable scaling behavior by observing more gains on larger networks. All this evidence reveals a promising future of generative pre-training on convnets. Codes and models are released at https://github.com/keyu-tian/SparK .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
- Gold-YOLO: Efficient Object Detector via Gather-and-Distribute MechanismChengcheng Wang, Wei He, Ying Nie, Jianyuan Guo 等NeurIPS 2023 · 被引用 732 次
- Architecture-Agnostic Masked Image Modeling - From ViT back to CNNSiyuan Li, Di Wu, Fang Wu, Zelin Zang 等ICML 2023 · 被引用 60 次
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksCanyu Zhao, Yanlong Sun, Mingyu Liu, Huanyi Zheng 等NeurIPS 2025 · 被引用 45 次
- DreamTeacher: Pretraining Image Backbones with Deep Generative ModelsDaiqing Li, Huan Ling, Amlan Kar, David Acuna 等ICCV 2023 · 被引用 37 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
相关 Paper
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang 等CVPR 2022 · 被引用 684 次
- Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object DetectionYuxin Fang, Shusheng Yang, Shijie Wang, Yixiao Ge 等ICCV 2023 · 被引用 67 次
- Masked Clustering Prediction for Unsupervised Point Cloud Pre-trainingBin Ren, Xiaoshui Huang, Mengyuan Liu, Hong Liu 等AAAI 2026 · 被引用 1 次
- Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingRenrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang 等NeurIPS 2022 · 被引用 445 次
- Take-A-Photo: 3D-to-2D Generative Pre-training of Point Cloud ModelsZiyi Wang, Xumin Yu, Yongming Rao, Jie Zhou 等ICCV 2023 · 被引用 34 次
