VaporTok: RL-Driven Adaptive Video Tokenizer with Prior & Task Awareness
Minghao Yang, Zechen Bai, Jing Lin, Haoqian Wang, Alex Jinpeng Wang
Abstract
Recent advances in visual tokenizers have demonstrated their effectiveness for multimodal large language models and autoregressive generative models. However, most existing visual tokenizers rely on a fixed downsampling rate at a given visual resolution, and consequently produce a constant number of visual tokens, ignoring the fact that visual information of varying complexity warrant different token budgets. Motivated by this observation, we propose an adaptive video tokenizer "VaporTok" with two core contributions: Probabilistic Taildrop : We introduce a novel taildrop mechanism that learns a truncation index sampling distribution conditioned on visual complexity of the video. During both training and inference, the decoder reconstructs videos at adaptive token lengths, allocating more tokens to complex videos and fewer to simpler ones. Parallel Sample GRPO with Vapor Reward : By leveraging the probability distribution produced by probabilistic taildrop, we reformulate the visual tokenization pipeline as a sequential decision process. To optimize this process, we propose a variant of GRPO and a composite reward encompassing token efficiency, reconstruction fidelity, and generative quality, thus enabling metrics-aware adaptive tokenization across diverse objectives. Extensive experiments on standard video generation benchmarks confirm our analysis, showing that our adaptive approach matches or outperforms fixed-rate baselines and naive taildrop while using fewer tokens.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e492074-afc1-407e-8dd9-06def18b65e6Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
Related papers
- VideoFlexTok: Flexible-Length Coarse-to-Fine Video TokenizationAndrei Atanov, Jesse Allardice, Roman Bachmann, Oğuzhan Fatih Kar et al.ICML 2026 · 3 citations
- EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive GenerationTianwei Xiong, Jun Hao Liew, Zilong Huang, Zhijie Lin et al.CVPR 2026 · 8 citations
- AdapTok: Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent SpaceYan Li, Changyao Tian, Renqiu Xia, Ning Liao et al.CVPR 2026
- ElasticTok: Adaptive Tokenization for Image and VideoWilson Yan, Volodymyr Mnih, Aleksandra Faust, Matei Zaharia et al.ICLR 2025
- TrajTok: Learning Trajectory Tokens Enhances Video UnderstandingChenhao Zheng, Jieyu Zhang, Jianing Zhang, Weikai Huang et al.CVPR 2026
