Efficient Vision-Language Pre-Training by Cluster Masking
Zihao Wei, Zixuan Pan, Andrew Owens
摘要
We propose a simple strategy for masking image patches during visual-language contrastive learning that improves the quality of the learned representations and the training speed. During each iteration of training, we randomly mask clusters of visually similar image patches, as measured by their raw pixel intensities. This provides an extra learning signal, beyond the contrastive training itself, since it forces a model to predict words for masked visual structures solely from context. It also speeds up training by reducing the amount of data used in each image. We evaluate the effectiveness of our model by pre-training on a number of bench-marks, finding that it outperforms other masking strategies, such as FLIP, on the quality of the learned representation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- WorDepth: Variational Language Prior for Monocular Depth EstimationZiyao Zeng, Daniel Wang, Fengyu Yang, Hyoungseob Park 等CVPR 2024 · 被引用 20 次
- SuperCLIP: CLIP with Simple Classification SupervisionWeiheng Zhao, Zilong Huang, Jiashi Feng, Xinggang WangNeurIPS 2025 · 被引用 6 次
- PowerCLIP: Powerset Alignment for Contrastive Pre-TrainingMasaki Kawamura, Nakamasa Inoue, Rintaro Yanagi, Hirokatsu Kataoka 等CVPR 2026 · 被引用 1 次
- Random Registers for Cross-Domain Few-Shot LearningShuai Yi, Yixiong Zou, Yuhua Li, Ruixuan LiICML 2025
- Self-guided Semantic Inspection for Zero-Shot Composed Image RetrievalJingjing Zhang, Lei Zhang, Zheren Fu, Bo Hu 等CVPR 2026
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- Scaling Language-Image Pre-Training via MaskingYanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer 等CVPR 2023
- Contextual Image Masking Modeling via Synergized Contrasting without View Augmentation for Faster and Better Visual PretrainingShaofeng Zhang, Feng Zhu, Rui Zhao, Junchi YanICLR 2023
- Learning Visual Representations via Language-Guided SamplingMohamed El Banani, Karan Desai, Justin JohnsonCVPR 2023
- One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual RepresentationXiaoyu Yang, Lijian Xu, Hongsheng Li, Shaoting ZhangICML 2025
- MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image PretrainingXiaoyi Dong, Jianmin Bao, Yinglin Zheng, Ting Zhang 等CVPR 2023
