Locally Hierarchical Auto-Regressive Modeling for Image Generation
Tackgeun You, Saehoon Kim, Chiheon Kim, Doyup Lee, Bohyung Han
摘要
We propose a locally hierarchical auto-regressive model with multiple resolutions of discrete codes. In the first stage of our algorithm, we represent an image with a pyramid of codes using Hierarchically Quantized Variational AutoEncoder (HQ-VAE), which disentangles the information contained in the multi-level codes. For an example of two-level codes, we create two separate pathways to carry high-level coarse structures of input images using top codes while compensating for missing fine details by constructing a residual connection for bottom codes. An appropriate selection of resizing operations for code embedding maps enables top codes to capture maximal information within images and the first stage algorithm achieves better performance on both vector quantization and image generation. The second stage adopts Hierarchically Quantized Transformer (HQ-Transformer) to process a sequence of local pyramids, which consist of a single top code and its corresponding bottom codes. Contrary to other hierarchical models, we sample bottom codes in parallel by exploiting the conditional independence assumption on the bottom codes. This assumption is naturally harvested from our first-stage model, HQ-VAE, where the bottom code learns to describe local details. On class-conditional and text-conditional generation benchmarks, our model shows competitive performance to previous AR models in terms of fidelity of generated images while enjoying lighter computational budgets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng 等ICML 2026 · 被引用 104 次
- Locality-Aware Generalizable Implicit Neural RepresentationDoyup Lee, Chiheon Kim, Minsu Cho, Wook-Shin HanNeurIPS 2023 · 被引用 25 次
- Kepler codebookJunrong Lian, Ziyue Dong, Pengxu Wei, Wei Ke 等ICML 2024 · 被引用 1 次
- MotionCtrl: A Real-Time Controllable Vision-Language-Motion ModelBin Cao, Sipeng Zheng, Ye Wang, Lujie Xia 等ICCV 2025 · 被引用 1 次
- Acquisition and Application of Novel Knowledge in Large Language ModelsZiyu Shang, Jianghan Liu, Zhizhao Luo, Peng Wang 等ACL 2025 · 被引用 1 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- Towards Accurate Image Coding: Improved Autoregressive Image Generation with Dynamic Vector QuantizationMengqi Huang, Zhendong Mao, Zhuowei Chen, Yongdong ZhangCVPR 2023
- Autoregressive Image Generation using Residual QuantizationDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho 等CVPR 2022 · 被引用 184 次
- Draft-and-Revise: Effective Image Generation with Contextual RQ-TransformerDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho 等NeurIPS 2022 · 被引用 36 次
- Hierarchical Sketch Induction for Paraphrase GenerationTom Hosking, Hao Tang, Mirella LapataACL 2022
- Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAEJialun Peng, Dong Liu, Songcen Xu, Houqiang LiCVPR 2021
