Pre-training with Random Orthogonal Projection Image Modeling
Maryam Haghighat, Peyman Moghadam, Shaheer Mohamed, Piotr Koniusz
摘要
Masked Image Modeling (MIM) is a powerful self-supervised strategy for visual pre-training without the use of labels. MIM applies random crops to input images, processes them with an encoder, and then recovers the masked inputs with a decoder, which encourages the network to capture and learn structural information about objects and scenes. The intermediate feature representations obtained from MIM are suitable for fine-tuning on downstream tasks. In this paper, we propose an Image Modeling framework based on random orthogonal projection instead of binary masking as in MIM. Our proposed Random Orthogonal Projection Image Modeling (ROPIM) reduces spatially-wise token information under guaranteed bound on the noise variance and can be considered as masking entire spatial image area under locally varying masking degrees. Since ROPIM uses a random subspace for the projection that realizes the masking step, the readily available complement of the subspace can be used during unmasking to promote recovery of removed information. In this paper, we show that using random orthogonal projection leads to superior performance compared to crop-based masking. We demonstrate state-of-the-art results on several popular benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Understanding and Mitigating Hyperbolic Dimensional Collapse in Graph Contrastive LearningYifei Zhang, Hao Zhu, Menglin Yang, Jiahong Liu 等KDD 2025 · 被引用 4 次
- Always Skip AttentionYiping Ji, Hemanth Saratchandran, Peyman Moghadam, Simon LuceyICCV 2025
- Robust Distillation via Untargeted and Targeted Intermediate Adversarial SamplesJunhao Dong, Piotr Koniusz, Junxi Chen, Z. Jane Wang 等CVPR 2024
- MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised LearningJingshan Hong, Haigen Hu, Huihuang Zhang, Qianwei Zhou 等AAAI 2026
- Bures-Isotropy Alignment: Manifold Learning of Generalized Category DiscoveryLuyao Tang, Kunze Huang, Chaoqi Chen, Cheng ChenICLR 2026
它引用的顶会 Paper32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- Correlational Image Modeling for Self-Supervised Visual Pre-TrainingWei Li, Jiahao Xie, Chen Change LoyCVPR 2023
- Architecture-Agnostic Masked Image Modeling - From ViT back to CNNSiyuan Li, Di Wu, Fang Wu, Zelin Zang 等ICML 2023 · 被引用 60 次
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier 等NeurIPS 2022 · 被引用 189 次
- Suppressing Non-Semantic Noise in Masked Image Modeling RepresentationsMartine Hjelkrem-Tan, Marius Aasan, Rwiddhi Chakraborty, Gabriel Y. Arteaga 等CVPR 2026
- Masked Frequency Modeling for Self-Supervised Visual Pre-TrainingJiahao Xie, Wei Li, Xiaohang Zhan, Ziwei Liu 等ICLR 2023 · 被引用 29 次
