Adaptive Length Image Tokenization via Recurrent Allocation
Shivam Duggal, Phillip Isola, Antonio Torralba, William T. Freeman
摘要
Current vision systems typically assign fixed-length representations to images, regardless of the information content. This contrasts with human intelligence -and even large language models-which allocate varying representational capacities based on entropy, context and familiarity. Inspired by this, we propose an approach to learn variable-length token representations for 2D images. Our encoder-decoder architecture recursively processes 2D image tokens, distilling them into 1D latent tokens over multiple iterations of recurrent rollouts. Each iteration refines the 2D tokens, updates the existing 1D latent tokens, and adaptively increases representational capacity by adding new tokens. This enables compression of images into a variable number of tokens, ranging from 32 to 256. We validate our tokenizer using reconstruction loss and FID metrics, demonstrating that token count aligns with image entropy, familiarity and downstream task requirements. Recurrent token processing with increasing representational capacity in each iteration shows signs of token specialization, revealing potential for object / part discovery. Code available at https://github.com/ ShivamDuggal4/adaptive-length-tokenizer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Dynamic Chunking for End-to-End Hierarchical Sequence ModelingSukjun Hwang, Brandon Wang, Albert GuICLR 2026 · 被引用 76 次
- Latent Denoising Makes Good TokenizersJiawei Yang, Tianhong Li, Lijie Fan, Yonglong Tian 等ICLR 2026 · 被引用 17 次
- CAT: Content-Adaptive Image TokenizationJunhong Shen, Kushal Tirumala, Michihiro Yasunaga, Ishan Misra 等NeurIPS 2025 · 被引用 17 次
- Single-pass Adaptive Image Tokenization for Minimum Program SearchShivam Duggal, Sanghyun Byun, Bill Freeman, Antonio Torralba 等NeurIPS 2025 · 被引用 11 次
- Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive GazingBaifeng Shi, Stephanie Fu, Long Lian, Hanrong Ye 等CVPR 2026 · 被引用 9 次
它引用的顶会 Paper22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
相关 Paper
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image GenerationTianwei Xiong, Jun Hao Liew, Zilong Huang, Jiashi Feng 等ICCV 2025 · 被引用 2 次
- Variable-Length Tokenization via Learnable Global Merging for Diffusion TransformersDong Hoon Lee, Seunghoon HongICML 2026
- VideoFlexTok: Flexible-Length Coarse-to-Fine Video TokenizationAndrei Atanov, Jesse Allardice, Roman Bachmann, Oğuzhan Fatih Kar 等ICML 2026 · 被引用 3 次
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual GenerationDivyansh Srivastava, Akshay Mehra, Pranav Maneriker, Debopam Sanyal 等CVPR 2026 · 被引用 1 次
- ElasticTok: Adaptive Tokenization for Image and VideoWilson Yan, Volodymyr Mnih, Aleksandra Faust, Matei Zaharia 等ICLR 2025
