Pre-training with Random Orthogonal Projection Image Modeling
Maryam Haghighat, Peyman Moghadam, Shaheer Mohamed, Piotr Koniusz
Abstract
Masked Image Modeling (MIM) is a powerful self-supervised strategy for visual pre-training without the use of labels. MIM applies random crops to input images, processes them with an encoder, and then recovers the masked inputs with a decoder, which encourages the network to capture and learn structural information about objects and scenes. The intermediate feature representations obtained from MIM are suitable for fine-tuning on downstream tasks. In this paper, we propose an Image Modeling framework based on random orthogonal projection instead of binary masking as in MIM. Our proposed Random Orthogonal Projection Image Modeling (ROPIM) reduces spatially-wise token information under guaranteed bound on the noise variance and can be considered as masking entire spatial image area under locally varying masking degrees. Since ROPIM uses a random subspace for the projection that realizes the masking step, the readily available complement of the subspace can be used during unmasking to promote recovery of removed information. In this paper, we show that using random orthogonal projection leads to superior performance compared to crop-based masking. We demonstrate state-of-the-art results on several popular benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c42b10e3-41d3-40ef-b540-00b02dc43abeCited by top-tier papers6
- Understanding and Mitigating Hyperbolic Dimensional Collapse in Graph Contrastive LearningYifei Zhang, Hao Zhu, Menglin Yang, Jiahong Liu et al.KDD 2025 · 4 citations
- Always Skip AttentionYiping Ji, Hemanth Saratchandran, Peyman Moghadam, Simon LuceyICCV 2025
- Robust Distillation via Untargeted and Targeted Intermediate Adversarial SamplesJunhao Dong, Piotr Koniusz, Junxi Chen, Z. Jane Wang et al.CVPR 2024
- MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised LearningJingshan Hong, Haigen Hu, Huihuang Zhang, Qianwei Zhou et al.AAAI 2026
- Bures-Isotropy Alignment: Manifold Learning of Generalized Category DiscoveryLuyao Tang, Kunze Huang, Chaoqi Chen, Cheng ChenICLR 2026
Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Correlational Image Modeling for Self-Supervised Visual Pre-TrainingWei Li, Jiahao Xie, Chen Change LoyCVPR 2023
- Architecture-Agnostic Masked Image Modeling - From ViT back to CNNSiyuan Li, Di Wu, Fang Wu, Zelin Zang et al.ICML 2023 · 60 citations
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier et al.NeurIPS 2022 · 189 citations
- Suppressing Non-Semantic Noise in Masked Image Modeling RepresentationsMartine Hjelkrem-Tan, Marius Aasan, Rwiddhi Chakraborty, Gabriel Y. Arteaga et al.CVPR 2026
- Masked Frequency Modeling for Self-Supervised Visual Pre-TrainingJiahao Xie, Wei Li, Xiaohang Zhan, Ziwei Liu et al.ICLR 2023 · 29 citations
