MCMAE: Masked Convolution Meets Masked Autoencoders
Peng Gao, Teli Ma, Hongsheng Li, Ziyi Lin, Jifeng Dai, Yu Qiao
摘要
Vision Transformers (ViT) become widely-adopted architectures for various vision tasks. Masked auto-encoding [2, 1, 28, 55] for feature pretraining and multi-scale hybrid convolution-transformer architectures [12, 21, 49, 34, 57] can further unleash the potentials of ViT, leading to state-of-the-art performances on image classification, detection and semantic segmentation. In this paper, our MCMAE framework demonstrates that multi-scale hybrid convolution-transformer can learn more discriminative representations via the mask auto-encoding scheme. However, directly using the original masking strategy leads to the heavy computational cost and pretraining-finetuning discrepancy. To tackle the issue, we adopt the masked convolution to prevent information leakage in the convolution blocks. A simple block-wise masking strategy is proposed to ensure computational efficiency. We also propose to more directly supervise the multi-scale features of the encoder to boost multi-scale features. MCMAE-Base improves ImageNet-1K finetuning accuracy by 1.4% compared with MAE-Base. On object detection, MCMAE-Base finetuned for only 25 epochs surpasses MAE-Base fined-tuned for 100 epochs by 2.9% AP box and 2.2% AP mask respectively. Code and pretrained models are available at https://github.com/Alpha-VL/ConvMAE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Hiera: A Hierarchical Vision Transformer without the Bells-and-WhistlesChaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei 等ICML 2023 · 被引用 388 次
- CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked AutoencodersAnthony Fuller, Koreen Millard, James R. GreenNeurIPS 2023 · 被引用 245 次
- Improving Pixel-based MIM by Reducing Wasted Modeling CapabilityYuan Liu, Songyang Zhang, Jiacheng Chen, Zhaohui Yu 等ICCV 2023 · 被引用 48 次
- TerraFM: A Scalable Foundation Model for Unified Multisensor Earth ObservationMuhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Muhammad Haris Khan 等ICLR 2026 · 被引用 30 次
- Synchronize Feature Extracting and Matching: A Single Branch Framework for 3D Object TrackingTeli Ma, Mengmeng Wang, Jimin Xiao, Huifeng Wu 等ICCV 2023 · 被引用 21 次
它引用的顶会 Paper32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object DetectionYuxin Fang, Shusheng Yang, Shijie Wang, Yixiao Ge 等ICCV 2023 · 被引用 67 次
- Masked Auto-Encoders Meet Generative Adversarial Networks and BeyondZhengcong Fei, Mingyuan Fan, Li Zhu, Junshi Huang 等CVPR 2023
- Asymmetric Masked Distillation for Pre-Training Small Foundation ModelsZhiyu Zhao, Bingkun Huang, Sen Xing, Gangshan Wu 等CVPR 2024 · 被引用 6 次
- MPViT: Multi-Path Vision Transformer for Dense PredictionYoungwan Lee, Jonghee Kim, Jeffrey Willette, Sung Ju HwangCVPR 2022 · 被引用 339 次
- MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision TransformersJihao Liu, Xin Huang, Jinliang Zheng, Yu Liu 等CVPR 2023
