The Image Local Autoregressive Transformer
Chenjie Cao, Yuxin Hong, Xiang Li, Chengrong Wang, Chengming Xu, Yanwei Fu, Xiangyang Xue
摘要
Recently, AutoRegressive (AR) models for the whole image generation empowered by transformers have achieved comparable or even better performance to Generative Adversarial Networks (GANs). Unfortunately, directly applying such AR models to edit/change local image regions, may suffer from the problems of missing global information, slow inference speed, and information leakage of local guidance. To address these limitations, we propose a novel model -- image Local Autoregressive Transformer (iLAT), to better facilitate the locally guided image synthesis. Our iLAT learns the novel local discrete representations, by the newly proposed local autoregressive (LA) transformer of the attention mask and convolution mechanism. Thus iLAT can efficiently synthesize the local image regions by key guidance information. Our iLAT is evaluated on various locally guided image syntheses, such as pose-guided person image synthesis and face editing. Both the quantitative and qualitative results show the efficacy of our model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- S2WAT: Image Style Transfer via Hierarchical Vision Transformer Using Strips Window AttentionChiyu Zhang, Xiaogang Xu, Lei Wang, Zaiyan Dai 等AAAI 2024 · 被引用 58 次
- Human MotionFormer: Transferring Human Motions with Vision TransformersHongyu Liu, Xintong Han, Chenbin Jin, Lihui Qian 等ICLR 2023 · 被引用 5 次
- What If: Understanding Motion Through Sparse InteractionsStefan Andreas Baumann, Nick Stracke, Timy Phan, Björn OmmerICCV 2025 · 被引用 3 次
- Krause Synchronization TransformersJingkun Liu, Yisong Yue, Max Welling, Yue SongICML 2026 · 被引用 1 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen 等ICCV 2019 · 被引用 1,990 次
相关 Paper
- Improving Visual Quality of Image Synthesis by A Token-based Generator with TransformersYanhong Zeng, Huan Yang, Hongyang Chao, Jianbo Wang 等NeurIPS 2021 · 被引用 31 次
- Deep Image Spatial Transformation for Person Image GenerationYurui Ren, Xiaoming Yu, Junming Chen, Thomas H. Li 等CVPR 2020
- Combining Attention with Flow for Person Image SynthesisYurui Ren, Yubo Wu, Thomas H. Li, Shan Liu 等ACM MM 2021 · 被引用 16 次
- A Latent Transformer for Disentangled Face Editing in Images and VideosXu Yao, Alasdair Newson, Yann Gousseau, Pierre HellierICCV 2021 · 被引用 97 次
- Holistic Tokenizer for Autoregressive Image GenerationAnlin Zheng, Haochen Wang, Yucheng Zhao, Weipeng Deng 等ICCV 2025 · 被引用 11 次
