CAT: Content-Adaptive Image Tokenization
Junhong Shen, Kushal Tirumala, Michihiro Yasunaga, Ishan Misra, Luke Zettlemoyer, Lili Yu, Chunting Zhou
Abstract
Most existing image tokenizers encode images into a fixed number of tokens or patches, overlooking the inherent variability in image complexity and introducing unnecessary computate overhead for simpler images. To address this, we propose Content-Adaptive Tokenizer (CAT), which dynamically adjusts representation capacity based on the image content and encodes simpler images into fewer tokens. We design (1) a caption-based evaluation system that leverages LLMs to predict content complexity and determine the optimal compression ratio for an image, and (2) a novel nested VAE architecture that performs variable-rate compression in a single model. Trained on images with varying complexity, CAT achieves an average of 15% reduction in rFID across seven detail-rich datasets containing text, humans, and complex textures. On natural image datasets like ImageNet and COCO, it reduces token usage by 18% while maintaining high-fidelity reconstructions. We further evaluate CAT on two downstream tasks. For image classification, CAT consistently improves top-1 accuracy across five datasets spanning diverse domains. For image generation, it boosts training throughput by 23% on ImageNet, leading to more efficient learning and improved FIDs over fixed-token baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f076eff9-13cc-4b62-b377-4d37fbd69dccCited by top-tier papers3
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual GenerationDivyansh Srivastava, Akshay Mehra, Pranav Maneriker, Debopam Sanyal et al.CVPR 2026 · 1 citation
- Content-Aware Dynamic Patchification for Efficient Video DiffusionSheng Li, Connelly Barnes, Mamshad Nayeem Rizve, Hongwu Peng et al.CVPR 2026
- AdapTok: Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent SpaceYan Li, Changyao Tian, Renqiu Xia, Ning Liao et al.CVPR 2026
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Adaptive Length Image Tokenization via Recurrent AllocationShivam Duggal, Phillip Isola, Antonio Torralba, William T. FreemanICLR 2025 · 1 citation
- SoftVQ-VAE: Efficient 1-Dimensional Continuous TokenizerHao Chen, Ze Wang, Xiang Li, Ximeng Sun et al.CVPR 2025
- CROP: Contextual Region-Oriented Visual Token PruningJiawei Guo, Feifei Zhai, Pu Jian, Qianrun Wei et al.EMNLP 2025
- Single-pass Adaptive Image Tokenization for Minimum Program SearchShivam Duggal, Sanghyun Byun, Bill Freeman, Antonio Torralba et al.NeurIPS 2025 · 11 citations
- BOLT: Fewer Tokens but More Performance Retention for Efficient Vision-Language Models InferenceJiahua Bao, Siyao Cheng, Jiaxing Du, Changjiang He et al.ACM MM 2025
