A Flat Vocabulary or a Rich Hierarchy? Re-introducing Intrinsic Structure Transforms the Autoregressive Image Generation
Lixuan He, Shikang Zheng
Abstract
Autoregressive (AR) models have shown great promise in image generation, yet they face a fundamental inefficiency stemming from their core component: a vast, unstructured vocabulary of visual tokens. By treating tokens as a flat set, standard models overlook the manifold structure where geometric proximity reflects semantic similarity. This oversight unnecessarily complicates the prediction task, hindering training efficiency and limiting generation quality. To resolve this, we propose Manifold-Aligned Semantic Clustering (MASC), a principled framework that constructs a hierarchical semantic tree directly from the codebook's intrinsic geometry. Utilizing a geometry-aware distance metric and density-driven agglomerative construction, MASC faithfully models the token embedding manifold. By transforming the flat, high-dimensional prediction into a structured hierarchical task, MASC introduces a powerful inductive bias that simplifies learning. Designed as a plug-and-play module, MASC accelerates training by up to 71% and significantly boosts generation quality, improving LlamaGen-XL's FID from 2.87 to 2.49. Crucially, MASC further serves as a convergence enabler for complex architectures. These results establish that structuring the prediction space is as vital as architectural innovation, elevating existing AR frameworks to state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
Related papers
- Improving Autoregressive Visual Generation with Cluster-Oriented Token PredictionTeng Hu, Jiangning Zhang, Ran Yi, Jieyu Weng et al.CVPR 2025
- Mirai: Autoregressive Visual Generation Needs ForesightYonghao Yu, Lang Huang, Zerun Wang, Runyi Li et al.CVPR 2026 · 1 citation
- Beyond Next-Token: Next-X Prediction for Autoregressive Visual GenerationSucheng Ren, Qihang Yu, Ju He, Xiaohui Shen et al.ICCV 2025 · 8 citations
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive GenerationXinyao Liao, QIYUAN HE, Kai Xu, Xiaoye Qu et al.CVPR 2026 · 6 citations
- Understand Before You Generate: Self-Guided Training for Autoregressive Image GenerationXiaoyu Yue, Zidong Wang, Yuqing Wang, Wenlong Zhang et al.NeurIPS 2025 · 9 citations
