Correlational Image Modeling for Self-Supervised Visual Pre-Training
Wei Li, Jiahao Xie, Chen Change Loy
Abstract
We introduce Correlational Image Modeling (CIM), a novel and surprisingly effective approach to self-supervised visual pre-training. Our CIM performs a simple pretext task: we randomly crop image regions (exemplars) from an input image (context) and predict correlation maps between the exemplars and the context. Three key designs enable correlational image modeling as a nontrivial and meaningful self-supervisory task. First, to generate useful exemplar-context pairs, we consider cropping image regions with various scales, shapes, rotations, and transformations. Second, we employ a bootstrap learning framework that involves online and target encoders. During pre-training, the former takes exemplars as inputs while the latter converts the context. Third, we model the output correlation maps via a simple cross-attention block, within which the context serves as queries and the exemplars offer values and keys. We show that CIM performs on par or better than the current state of the art on self-supervised and transfer benchmarks. Code is available at https://github.com/weivision/Correlational-Image-Modeling.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6510ca4d-e6c9-4480-ada5-8c6a63fddee6Cited by top-tier papers7
- Pre-training with Random Orthogonal Projection Image ModelingMaryam Haghighat, Peyman Moghadam, Shaheer Mohamed, Piotr KoniuszICLR 2024 · 15 citations
- SCE-MAE: Selective Correspondence Enhancement with Masked Autoencoder for Self-Supervised Landmark EstimationKejia Yin, Varshanth S. Rao, Ruowei Jiang, Xudong Liu et al.CVPR 2024 · 1 citation
- Frequency-Guided Masking for Enhanced Vision Self-Supervised LearningAmin Karimi Monsefi, Mengxi Zhou, Nastaran Karimi Monsefi, Ser-Nam Lim et al.ICLR 2025
- OMG-Seg: Is One Model Good Enough for all Segmentation?Xiangtai Li, Haobo Yuan, Wei Li, Henghui Ding et al.CVPR 2024
- Test-Time Visual In-Context TuningJiahao Xie, Alessio Tonioni, Nathalie Rauschmayr, Federico Tombari et al.CVPR 2025
Builds on42
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Corrupted Image Modeling for Self-Supervised Visual Pre-TrainingYuxin Fang, Li Dong, Hangbo Bao, Xinggang Wang et al.ICLR 2023 · 23 citations
- Exploring Stochastic Autoregressive Image Modeling for Visual RepresentationYu Qi, Fan Yang, Yousong Zhu, Yufei Liu et al.AAAI 2023 · 18 citations
- Architecture-Agnostic Masked Image Modeling - From ViT back to CNNSiyuan Li, Di Wu, Fang Wu, Zelin Zang et al.ICML 2023 · 60 citations
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier et al.NeurIPS 2022 · 189 citations
- Self-Supervised Learning of Pretext-Invariant RepresentationsIshan Misra, Laurens van der MaatenCVPR 2020
