OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
Junke Wang, Yi Jiang, Zehuan Yuan, Bingyue Peng, Zuxuan Wu, Yu-Gang Jiang
摘要
Tokenizer, serving as a translator to map the intricate visual data into a compact latent space, lies at the core of visual generative models. Based on the finding that existing tokenizers are tailored to image or video inputs, this paper presents OmniTokenizer, a transformer-based tokenizer for joint image and video tokenization. OmniTokenizer is designed with a spatial-temporal decoupled architecture, which integrates window and causal attention for spatial and temporal modeling. To exploit the complementary nature of image and video data, we further propose a progressive training strategy, where OmniTokenizer is first trained on image data on a fixed resolution to develop the spatial encoding capacity and then jointly trained on image and video data on multiple resolutions to learn the temporal dynamics. OmniTokenizer, for the first time, handles both image and video inputs within a unified framework and proves the possibility of realizing their synergy. Extensive experiments demonstrate that OmniTokenizer achieves state-of-the-art (SOTA) reconstruction performance on various image and video datasets, e.g., 1.11 reconstruction FID on ImageNet and 42 reconstruction FVD on UCF-101, beating the previous SOTA methods by 13% and 26%, respectively. Additionally, we also show that when integrated with OmniTokenizer, both language model-based approaches and diffusion models can realize advanced visual synthesis performance, underscoring the superiority and versatility of our method. Code is available at https://github.com/FoundationVision/OmniTokenizer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual GenerationJinlai Liu, Jian Han, Bin Yan, Hui Wu 等NeurIPS 2025 · 被引用 45 次
- Generative Pre-trained Autoregressive Diffusion TransformerYuan Zhang, Jiacheng Jiang, Guoqing Ma, Zhiying Lu 等NeurIPS 2025 · 被引用 19 次
- CODA: Repurposing Continuous VAEs for Discrete TokenizationZeyu Liu, Zanlin Ni, Yeguo Hua, Xin Deng 等ICCV 2025 · 被引用 9 次
- Video-GPT via Next Clip DiffusionShaobin Zhuang, Zhipeng Huang, Ying Zhang, Fangyikang Wang 等ICLR 2026 · 被引用 9 次
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual ReconstructionShaobin Zhuang, Yiwei Guo, Fangyikang Wang, Canmiao Fu 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- AToken: A Unified Tokenizer for VisionJiasen Lu, Liangchen Song, Mingze Xu, Byeongjoo Ahn 等CVPR 2026 · 被引用 33 次
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari 等ICLR 2024 · 被引用 609 次
- Efficient Long Video Tokenization via Coordinate-based Patch ReconstructionHuiwon Jang, Sihyun Yu, Jinwoo Shin, Pieter Abbeel 等CVPR 2025
- VideoMAETok: Boosting Video Diffusion Models via Masked Autoencoders as TokenizersZhan Tong, Tinne TuytelaarsICML 2026
- End-to-End Autoregressive Image Generation with 1D Semantic TokenizerWenda Chu, Bingliang Zhang, Jiaqi Han, Yizhuo Li 等ICML 2026 · 被引用 2 次
