On the Role of Discrete Tokenization in Visual Representation Learning
Tianqi Du, Yifei Wang, Yisen Wang
Abstract
In the realm of self-supervised learning (SSL), masked image modeling (MIM) has gained popularity alongside contrastive learning methods. MIM involves reconstructing masked regions of input images using their unmasked portions. A notable subset of MIM methodologies employs discrete tokens as the reconstruction target, but the theoretical underpinnings of this choice remain underexplored. In this paper, we explore the role of these discrete tokens, aiming to unravel their benefits and limitations. Building upon the connection between MIM and contrastive learning, we provide a comprehensive theoretical understanding on how discrete tokenization affects the model's generalization capabilities. Furthermore, we propose a novel metric named TCAS, which is specifically designed to assess the effectiveness of discrete tokens within the MIM framework. Inspired by this metric, we contribute an innovative tokenizer design and propose a corresponding MIM method named ClusterMIM. It demonstrates superior performance on a variety of benchmark datasets and ViT backbones. Code is available at https://github.com/PKU-ML/ClusterMIM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- VTD-CLIP: Video-to-Text Discretization via Prompting CLIPWencheng Zhu, Yuexin Wang, Hongxuan Li, Pengfei ZhuAAAI 2026 · 2 citations
- Multimodal Medical Code TokenizerXiaorui Su, Shvat Messica, Yepeng Huang, Ruth Johnson et al.ICML 2025 · 2 citations
- Closest Neighbors are Harmful for Lightweight Masked Auto-encodersJian Meng, Ahmed Hassan, Li Yang, Deliang Fan et al.CVPR 2025
- Identifying and Understanding Cross-Class Features in Adversarial TrainingZeming Wei, Steven Y. Guo, Yisen WangICML 2025
- Projection Head is Secretly an Information BottleneckZhuo Ouyang, Kaiwen Hu, Qi Zhang, Yifei Wang et al.ICLR 2025
Builds on24
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- What Do Self-Supervised Vision Transformers Learn?Namuk Park, Wonjae Kim, Byeongho Heo, Taekyung Kim et al.ICLR 2023 · 16 citations
- Masked Image Modeling with Denoising ContrastKun Yi, Yixiao Ge, Xiaotong Li, Shusheng Yang et al.ICLR 2023 · 8 citations
- MimCo: Masked Image Modeling Pre-training with Contrastive TeacherQiang Zhou, Chaohui Yu, Hao Luo, Zhibin Wang et al.ACM MM 2022 · 16 citations
- Hybrid Distillation: Connecting Masked Autoencoders with Contrastive LearnersBowen Shi, Xiaopeng Zhang, Yaoming Wang, Jin Li et al.ICLR 2024 · 10 citations
- Understanding Masked Image Modeling via Learning Occlusion Invariant FeatureXiangwen Kong, Xiangyu ZhangCVPR 2023
