Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction
Teng Hu, Jiangning Zhang, Ran Yi, Jieyu Weng, Yabiao Wang, Xianfang Zeng, Zhucun Xue, Lizhuang Ma
摘要
Employing LLMs for visual generation has recently become a research focus. However, the existing methods primarily transfer the LLM architecture to visual generation but rarely investigate the fundamental differences between language and vision. This oversight may lead to suboptimal utilization of visual generation capabilities within the LLM framework. In this paper, we explore the characteristics of visual embedding space under the LLM framework and discover that the correlation between visual embeddings can help achieve more stable and robust generation results. We present IAR, an Improved AutoRegressive Visual Generation Method that enhances the training efficiency and generation quality of LLM-based visual generation models. Firstly, we propose a Codebook Rearrangement strategy that uses balanced k-means clustering algorithm to rearrange the visual codebook into clusters, ensuring high similarity among visual features within each cluster. Leveraging the rearranged codebook, we propose a Cluster-oriented Cross-entropy Loss that guides the model to correctly predict the cluster where the target token is located. This approach ensures that even if the model predicts the wrong token index, there is a high probability the predicted token is located in the correct cluster, which significantly enhances the generation quality and robustness. Extensive experiments demonstrate that our IAR consistently enhances the model training efficiency and performance from 100M to 1.4B, reducing the training time by half while achieving the same FID. Additionally, IAR can be applied to various LLM-based visual generation models and adheres to the scaling law, providing a promising direction for future research in LLM-based visual generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Harmony: Harmonizing Audio and Video Generation through Cross-Task SynergyTeng Hu, Zhentao Yu, Guozhen Zhang, Zihan Su 等CVPR 2026 · 被引用 21 次
- Hierarchical Image Tokenization for Multi-Scale Image Super ResolutionIsma Hadji, Enrique Sanchez, Adrian Bulat, Brais Martinez 等ICML 2026
- Image Token Matters: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent EditingWeixing Wang, Zifeng Ding, Jindong Gu, Rui Cao 等NeurIPS 2025
- A Flat Vocabulary or a Rich Hierarchy? Re-introducing Intrinsic Structure Transforms the Autoregressive Image GenerationLixuan He, Shikang ZhengICML 2026
- Pinco: Position-Induced Consistent Adapter for Diffusion Transformer in Foreground-Conditioned InpaintingGuangben Lu, Yuzhen Du, Yizhe Tang, Zhimin Sun 等ICCV 2025
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Elucidating the design space of language models for image generationXuantong Liu, Shaozhe Hao, Xianbiao Qi, Tianyang Hu 等ICML 2025
- Mirai: Autoregressive Visual Generation Needs ForesightYonghao Yu, Lang Huang, Zerun Wang, Runyi Li 等CVPR 2026 · 被引用 1 次
- Randomized Autoregressive Visual GenerationQihang Yu, Ju He, Xueqing Deng, Xiaohui Shen 等ICCV 2025 · 被引用 10 次
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive GenerationXinyao Liao, QIYUAN HE, Kai Xu, Xiaoye Qu 等CVPR 2026 · 被引用 6 次
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
