Holistic Tokenizer for Autoregressive Image Generation
Anlin Zheng, Haochen Wang, Yucheng Zhao, Weipeng Deng, Tiancai Wang, Xiangyu Zhang, Xiaojuan Qi
Abstract
Vanilla autoregressive image generation models generate visual tokens step-by-step, limiting their ability to capture holistic relationships among token sequences. Moreover, because most visual tokenizers map local image patches into latent tokens, global information is limited. To address this, we introduce Hita, a novel image tokenizer for autoregressive () image generation. It introduces a holistic-to-local tokenization scheme with learnable holistic queries and local patch tokens. Hita incorporates two key strategies to better align with the AR generation process: 1) arranging a sequential structure with holistic tokens at the beginning, followed by patch-level tokens, and using causal attention to maintain awareness of previous tokens; and 2) adopting a lightweight fusion module before feeding the de-quantized tokens into the decoder to control information flow and prioritize holistic tokens. Extensive experiments show that Hita accelerates the training speed of AR generators and outperforms those trained with vanilla tokenizers, achieving 2.59 FID and 281.9 IS on the ImageNet benchmark. Detailed analysis of the holistic representation highlights its ability to capture global image properties, such as textures, materials, and shapes. Additionally, Hita also demonstrates effectiveness in zero-shot style transfer and image in-painting. The code is available at https://github.com/CVMI-Lab/Hita.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0e13201-3b32-49b4-91f3-352673b8be8eCited by top-tier papers2
- From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image FusionYuchen Xian, Yunqiu Xu, Yang He, Yi YangICML 2026 · 2 citations
- Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font GenerationHaonan Cai, Yuxuan Luo, Zhouhui LianCVPR 2026 · 1 citation
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive ModelPingyu Wu, Kai Zhu, Yu Liu, Longxiang Tang et al.ICLR 2026 · 16 citations
- LARP: Tokenizing Videos with a Learned Autoregressive Generative PriorHanyu Wang, Saksham Suri, Yixuan Ren, Hao Chen et al.ICLR 2025
- GloTok: Global Perspective Tokenizer for Image Reconstruction and GenerationXuan Zhao, Zhongyu Zhang, Yuge Huang, Yuxi Mi et al.AAAI 2026 · 1 citation
- reAR: Rethinking Visual Autoregressive Models via Token-wise Consistency RegularizationQiyuan He, Yicong Li, Haotian Ye, Jinghao Wang et al.ICLR 2026 · 4 citations
- Next Patch Prediction for AutoRegressive Visual GenerationYatian Pang, Peng Jin, Shuo Yang, Bin Zhu et al.AAAI 2026 · 24 citations
