The Image Local Autoregressive Transformer
Chenjie Cao, Yuxin Hong, Xiang Li, Chengrong Wang, Chengming Xu, Yanwei Fu, Xiangyang Xue
Abstract
Recently, AutoRegressive (AR) models for the whole image generation empowered by transformers have achieved comparable or even better performance to Generative Adversarial Networks (GANs). Unfortunately, directly applying such AR models to edit/change local image regions, may suffer from the problems of missing global information, slow inference speed, and information leakage of local guidance. To address these limitations, we propose a novel model -- image Local Autoregressive Transformer (iLAT), to better facilitate the locally guided image synthesis. Our iLAT learns the novel local discrete representations, by the newly proposed local autoregressive (LA) transformer of the attention mask and convolution mechanism. Thus iLAT can efficiently synthesize the local image regions by key guidance information. Our iLAT is evaluated on various locally guided image syntheses, such as pose-guided person image synthesis and face editing. Both the quantitative and qualitative results show the efficacy of our model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69b16f6b-b555-4e8b-84c5-38795f99a7f9Cited by top-tier papers4
- S2WAT: Image Style Transfer via Hierarchical Vision Transformer Using Strips Window AttentionChiyu Zhang, Xiaogang Xu, Lei Wang, Zaiyan Dai et al.AAAI 2024 · 58 citations
- Human MotionFormer: Transferring Human Motions with Vision TransformersHongyu Liu, Xintong Han, Chenbin Jin, Lihui Qian et al.ICLR 2023 · 5 citations
- What If: Understanding Motion Through Sparse InteractionsStefan Andreas Baumann, Nick Stracke, Timy Phan, Björn OmmerICCV 2025 · 3 citations
- Krause Synchronization TransformersJingkun Liu, Yisong Yue, Max Welling, Yue SongICML 2026 · 1 citation
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
Related papers
- Improving Visual Quality of Image Synthesis by A Token-based Generator with TransformersYanhong Zeng, Huan Yang, Hongyang Chao, Jianbo Wang et al.NeurIPS 2021 · 31 citations
- Deep Image Spatial Transformation for Person Image GenerationYurui Ren, Xiaoming Yu, Junming Chen, Thomas H. Li et al.CVPR 2020
- Combining Attention with Flow for Person Image SynthesisYurui Ren, Yubo Wu, Thomas H. Li, Shan Liu et al.ACM MM 2021 · 16 citations
- A Latent Transformer for Disentangled Face Editing in Images and VideosXu Yao, Alasdair Newson, Yann Gousseau, Pierre HellierICCV 2021 · 97 citations
- Holistic Tokenizer for Autoregressive Image GenerationAnlin Zheng, Haochen Wang, Yucheng Zhao, Weipeng Deng et al.ICCV 2025 · 11 citations
