ElasticTok: Adaptive Tokenization for Image and Video
Wilson Yan, Volodymyr Mnih, Aleksandra Faust, Matei Zaharia, Pieter Abbeel, Hao Liu
摘要
Efficient video tokenization remains a key bottleneck in learning general purpose vision models that are capable of processing long video sequences. Prevailing approaches are restricted to encoding videos to a fixed number of tokens, where too few tokens will result in overly lossy encodings, and too many tokens will result in prohibitively long sequence lengths. In this work, we introduce ElasticTok, a method that conditions on prior frames to adaptively encode a frame into a variable number of tokens. To enable this in a computationally scalable way, we propose a masking technique that drops a random number of tokens at the end of each frames's token encoding. During inference, ElasticTok can dynamically allocate tokens when needed -more complex data can leverage more tokens, while simpler data only needs a few tokens. Our empirical evaluations on images and video demonstrate the effectiveness of our approach in efficient token usage, paving the way for future development of more powerful multimodal models, world models, and agents. Video examples of using ElasticTok can be found on our website: largeworldmodel.github.io/elastictok 0 40 144 228 300 343 408 491 GT Recon t=0s t=21s Time ⋆ To whom correspondence should be addressed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- AToken: A Unified Tokenizer for VisionJiasen Lu, Liangchen Song, Mingze Xu, Byeongjoo Ahn 等CVPR 2026 · 被引用 33 次
- Latent Denoising Makes Good TokenizersJiawei Yang, Tianhong Li, Lijie Fan, Yonglong Tian 等ICLR 2026 · 被引用 17 次
- CAT: Content-Adaptive Image TokenizationJunhong Shen, Kushal Tirumala, Michihiro Yasunaga, Ishan Misra 等NeurIPS 2025 · 被引用 17 次
- Adapting Self-Supervised Representations as a Latent Space for Efficient GenerationMing Gui, Johannes Schusterbauer, Timy Phan, Felix Krause 等ICLR 2026 · 被引用 13 次
- Single-pass Adaptive Image Tokenization for Minimum Program SearchShivam Duggal, Sanghyun Byun, Bill Freeman, Antonio Torralba 等NeurIPS 2025 · 被引用 11 次
它引用的顶会 Paper22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
相关 Paper
- VideoFlexTok: Flexible-Length Coarse-to-Fine Video TokenizationAndrei Atanov, Jesse Allardice, Roman Bachmann, Oğuzhan Fatih Kar 等ICML 2026 · 被引用 3 次
- AdapTok: Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent SpaceYan Li, Changyao Tian, Renqiu Xia, Ning Liao 等CVPR 2026
- EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive GenerationTianwei Xiong, Jun Hao Liew, Zilong Huang, Zhijie Lin 等CVPR 2026 · 被引用 8 次
- VaporTok: RL-Driven Adaptive Video Tokenizer with Prior & Task AwarenessMinghao Yang, Zechen Bai, Jing Lin, Haoqian Wang 等NeurIPS 2025 · 被引用 1 次
- Efficient Long Video Tokenization via Coordinate-based Patch ReconstructionHuiwon Jang, Sihyun Yu, Jinwoo Shin, Pieter Abbeel 等CVPR 2025
