InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
Haotian Ye, Qiyuan He, Jiaqi Han, Puheng Li, Jiaojiao Fan, Zekun Hao, Fitsum Reda, Yogesh Balaji, Huayu Chen, Sheng Liu, Angela Yao, James Y. Zou
Abstract
Accurate and efficient discrete video tokenization is essential for long video sequences processing. Yet, the inherent complexity and variable information density of videos present a significant bottleneck for current tokenizers, which rigidly compress all content at a fixed rate, leading to redundancy or information loss. Drawing inspiration from Shannon's information theory, this paper introduces , a principled framework for adaptive video tokenization. We rigorously prove that existing data-agnostic training methods are suboptimal in representation length, and present a novel evidence lower bound (ELBO)-based algorithm that approaches theoretical optimality. Leveraging this framework, we develop a transformer-based adaptive compressor that enables adaptive tokenization. Empirical results demonstrate state-of-the-art compression performance, saving tokens without influence on performance, and achieving compression rates while still outperforming prior heuristic adaptive approaches. By allocating tokens according to informational richness, enables a more compressed yet accurate tokenization for video representation, offering valuable insights for future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari et al.ICLR 2024 · 609 citations
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 442 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
- An Image is Worth 32 Tokens for Reconstruction and GenerationQihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2024 · 331 citations
Related papers
- UniComp: Rethinking Video Compression Through Informational UniquenessChao Yuan, Shimin Chen, Minliang Lin, Limeng Qiao et al.CVPR 2026 · 6 citations
- VQToken: Neural Discrete Token Representation Learning for Extreme Token Reduction in Video Large Language ModelsHaichao Zhang, Yun FuNeurIPS 2025 · 15 citations
- Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language ModelsYucheng Zhou, Jihai Zhang, Guanjie Chen, Jianbing Shen et al.AAAI 2026
- Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token CompressorsXiangchen Wang, Jinrui Zhang, Teng Wang, Haigang Zhang et al.EMNLP 2025 · 2 citations
- ElasticTok: Adaptive Tokenization for Image and VideoWilson Yan, Volodymyr Mnih, Aleksandra Faust, Matei Zaharia et al.ICLR 2025
