Masked Frequency Modeling for Self-Supervised Visual Pre-Training
Jiahao Xie, Wei Li, Xiaohang Zhan, Ziwei Liu, Yew-Soon Ong, Chen Change Loy
摘要
We present Masked Frequency Modeling (MFM), a unified frequency-domainbased approach for self-supervised pre-training of visual models. Instead of randomly inserting mask tokens to the input embeddings in the spatial domain, in this paper, we shift the perspective to the frequency domain. Specifically, MFM first masks out a portion of frequency components of the input image and then predicts the missing frequencies on the frequency spectrum. Our key insight is that predicting masked components in the frequency domain is more ideal to reveal underlying image patterns rather than predicting masked patches in the spatial domain, due to the heavy spatial redundancy. Our findings suggest that with the right configuration of mask-and-predict strategy, both the structural information within high-frequency components and the low-level statistics among low-frequency counterparts are useful in learning good representations. For the first time, MFM demonstrates that, for both ViT and CNN, a simple non-Siamese framework can learn meaningful representations even using none of the following: (i) extra data, (ii) extra model, (iii) mask token. Experimental results on image classification and semantic segmentation, as well as several robustness benchmarks show the competitive performance and advanced robustness of MFM compared with recent masked image modeling approaches. Furthermore, we also comprehensively investigate the effectiveness of classical image restoration tasks for representation learning from a unified frequency perspective and reveal their intriguing relations with our MFM approach. Project page: https://www.mmlab-ntu.com/project/mfm/index.html .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Frequency Guidance Matters in Few-Shot LearningHao Cheng, Siyuan Yang, Joey Tianyi Zhou, Lanqing Guo 等ICCV 2023 · 被引用 48 次
- FreeKD: Knowledge Distillation via Semantic Frequency PromptYuan Zhang, Tao Huang, Jiaming Liu, Tao Jiang 等CVPR 2024 · 被引用 26 次
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 被引用 18 次
- Pre-training with Random Orthogonal Projection Image ModelingMaryam Haghighat, Peyman Moghadam, Shaheer Mohamed, Piotr KoniuszICLR 2024 · 被引用 15 次
- Efficient Vision-Language Pre-Training by Cluster MaskingZihao Wei, Zixuan Pan, Andrew OwensCVPR 2024 · 被引用 4 次
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
相关 Paper
- The Devil Is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-trainingHao Liu, Xinghua Jiang, Xin Li, Antai Guo 等AAAI 2023 · 被引用 45 次
- Good Helper Is around You: Attention-Driven Masked Image ModelingZhengqi Liu, Jie Gui, Hao LuoAAAI 2023 · 被引用 36 次
- Architecture-Agnostic Masked Image Modeling - From ViT back to CNNSiyuan Li, Di Wu, Fang Wu, Zelin Zang 等ICML 2023 · 被引用 60 次
- MAGE: MAsked Generative Encoder to Unify Representation Learning and Image SynthesisTianhong Li, Huiwen Chang, Shlok Kumar Mishra, Han Zhang 等CVPR 2023
- DDAE: Towards Deep Dynamic Vision BERT PretrainingHonghao Chen, Xiangwen Kong, Xiangyu Zhang, Xin Zhao 等AAAI 2024 · 被引用 1 次
