Locally Hierarchical Auto-Regressive Modeling for Image Generation
Tackgeun You, Saehoon Kim, Chiheon Kim, Doyup Lee, Bohyung Han
Abstract
We propose a locally hierarchical auto-regressive model with multiple resolutions of discrete codes. In the first stage of our algorithm, we represent an image with a pyramid of codes using Hierarchically Quantized Variational AutoEncoder (HQ-VAE), which disentangles the information contained in the multi-level codes. For an example of two-level codes, we create two separate pathways to carry high-level coarse structures of input images using top codes while compensating for missing fine details by constructing a residual connection for bottom codes. An appropriate selection of resizing operations for code embedding maps enables top codes to capture maximal information within images and the first stage algorithm achieves better performance on both vector quantization and image generation. The second stage adopts Hierarchically Quantized Transformer (HQ-Transformer) to process a sequence of local pyramids, which consist of a single top code and its corresponding bottom codes. Contrary to other hierarchical models, we sample bottom codes in parallel by exploiting the conditional independence assumption on the bottom codes. This assumption is naturally harvested from our first-stage model, HQ-VAE, where the bottom code learns to describe local details. On class-conditional and text-conditional generation benchmarks, our model shows competitive performance to previous AR models in terms of fidelity of generated images while enjoying lighter computational budgets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 211f930d-a15c-4f28-b3bc-8fbd02e8da40Cited by top-tier papers7
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng et al.ICML 2026 · 104 citations
- Locality-Aware Generalizable Implicit Neural RepresentationDoyup Lee, Chiheon Kim, Minsu Cho, Wook-Shin HanNeurIPS 2023 · 25 citations
- Kepler codebookJunrong Lian, Ziyue Dong, Pengxu Wei, Wei Ke et al.ICML 2024 · 1 citation
- MotionCtrl: A Real-Time Controllable Vision-Language-Motion ModelBin Cao, Sipeng Zheng, Ye Wang, Lujie Xia et al.ICCV 2025 · 1 citation
- Acquisition and Application of Novel Knowledge in Large Language ModelsZiyu Shang, Jianghan Liu, Zhizhao Luo, Peng Wang et al.ACL 2025 · 1 citation
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Towards Accurate Image Coding: Improved Autoregressive Image Generation with Dynamic Vector QuantizationMengqi Huang, Zhendong Mao, Zhuowei Chen, Yongdong ZhangCVPR 2023
- Autoregressive Image Generation using Residual QuantizationDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho et al.CVPR 2022 · 184 citations
- Draft-and-Revise: Effective Image Generation with Contextual RQ-TransformerDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho et al.NeurIPS 2022 · 36 citations
- Hierarchical Sketch Induction for Paraphrase GenerationTom Hosking, Hao Tang, Mirella LapataACL 2022
- Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAEJialun Peng, Dong Liu, Songcen Xu, Houqiang LiCVPR 2021
