The Devil Is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-training
Hao Liu, Xinghua Jiang, Xin Li, Antai Guo, Yiqing Hu, Deqiang Jiang, Bo Ren
Abstract
The self-supervised Masked Image Modeling (MIM) schema, following "mask-and-reconstruct" pipeline of recovering contents from masked image, has recently captured the increasing interest in the community, owing to the excellent ability of learning visual representation from unlabeled data. Aiming at learning representations with high semantics abstracted, a group of works attempts to reconstruct non-semantic pixels with large-ratio masking strategy, which may suffer from "over-smoothing" problem, while others directly infuse semantics into targets in offline way requiring extra data. Different from them, we shift the perspective to the Fourier domain which naturally has global perspective and present a new Masked Image Modeling (MIM), termed Geminated Gestalt Autoencoder (Ge 2 -AE) for visual pre-training. Specifically, we equip our model with geminated decoders in charge of reconstructing image contents from both pixel and frequency space, where each other serves as not only the complementation but also the reciprocal constraints. Through this way, more robust representations can be learned in the pre-trained encoders, of which the effectiveness is confirmed by the juxtaposing experimental results on downstream recognition tasks. We also conduct several quantitative and qualitative experiments to investigate the learning behavior of our method. To our best knowledge, this is the first MIM work to solve the visual pre-training through the lens of frequency domain. * Equal contribution. † Contact person.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Masked Frequency Modeling for Self-Supervised Visual Pre-TrainingJiahao Xie, Wei Li, Xiaohang Zhan, Ziwei Liu et al.ICLR 2023 · 29 citations
- DiMSUM: Diffusion Mamba - A Scalable and Unified Spatial-Frequency Method for Image GenerationHao Phung, Quan Dao, Trung Tuan Dao, Viet Hoang Phan et al.NeurIPS 2024 · 21 citations
- ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion ProcessChangyao Tian, Chenxin Tao, Jifeng Dai, Hao Li et al.ICLR 2024 · 19 citations
- RevColV2: Exploring Disentangled Representations in Masked Image ModelingQi Han, Yuxuan Cai, Xiangyu ZhangNeurIPS 2023 · 16 citations
- RoMA: Scaling up Mamba-based Foundation Models for Remote SensingFengxiang Wang, Yulin Wang, Mingshuo Chen, Haotian Wang et al.NeurIPS 2025 · 14 citations
Builds on29
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Good Helper Is around You: Attention-Driven Masked Image ModelingZhengqi Liu, Jie Gui, Hao LuoAAAI 2023 · 36 citations
- MimCo: Masked Image Modeling Pre-training with Contrastive TeacherQiang Zhou, Chaohui Yu, Hao Luo, Zhibin Wang et al.ACM MM 2022 · 16 citations
- Global Patch-wise Attention is Masterful Facilitator for Masked Image ModelingGongli Xi, Ye Tian, Mengyu Yang, Lanshan Zhang et al.ACM MM 2024 · 1 citation
- MAGE: MAsked Generative Encoder to Unify Representation Learning and Image SynthesisTianhong Li, Huiwen Chang, Shlok Kumar Mishra, Han Zhang et al.CVPR 2023
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier et al.NeurIPS 2022 · 189 citations
