Efficient Vision-Language Pre-Training by Cluster Masking
Zihao Wei, Zixuan Pan, Andrew Owens
Abstract
We propose a simple strategy for masking image patches during visual-language contrastive learning that improves the quality of the learned representations and the training speed. During each iteration of training, we randomly mask clusters of visually similar image patches, as measured by their raw pixel intensities. This provides an extra learning signal, beyond the contrastive training itself, since it forces a model to predict words for masked visual structures solely from context. It also speeds up training by reducing the amount of data used in each image. We evaluate the effectiveness of our model by pre-training on a number of bench-marks, finding that it outperforms other masking strategies, such as FLIP, on the quality of the learned representation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5120027f-a732-4dbc-90af-ad33ffd51f3aCited by top-tier papers9
- WorDepth: Variational Language Prior for Monocular Depth EstimationZiyao Zeng, Daniel Wang, Fengyu Yang, Hyoungseob Park et al.CVPR 2024 · 20 citations
- SuperCLIP: CLIP with Simple Classification SupervisionWeiheng Zhao, Zilong Huang, Jiashi Feng, Xinggang WangNeurIPS 2025 · 6 citations
- PowerCLIP: Powerset Alignment for Contrastive Pre-TrainingMasaki Kawamura, Nakamasa Inoue, Rintaro Yanagi, Hirokatsu Kataoka et al.CVPR 2026 · 1 citation
- Random Registers for Cross-Domain Few-Shot LearningShuai Yi, Yixiong Zou, Yuhua Li, Ruixuan LiICML 2025
- Self-guided Semantic Inspection for Zero-Shot Composed Image RetrievalJingjing Zhang, Lei Zhang, Zheren Fu, Bo Hu et al.CVPR 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Scaling Language-Image Pre-Training via MaskingYanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer et al.CVPR 2023
- Contextual Image Masking Modeling via Synergized Contrasting without View Augmentation for Faster and Better Visual PretrainingShaofeng Zhang, Feng Zhu, Rui Zhao, Junchi YanICLR 2023
- Learning Visual Representations via Language-Guided SamplingMohamed El Banani, Karan Desai, Justin JohnsonCVPR 2023
- One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual RepresentationXiaoyu Yang, Lijian Xu, Hongsheng Li, Shaoting ZhangICML 2025
- MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image PretrainingXiaoyi Dong, Jianmin Bao, Yinglin Zheng, Ting Zhang et al.CVPR 2023
