Taming Transformers for High-Resolution Image Synthesis
Patrick Esser, Robin Rombach, Björn Ommer
摘要
Designed to learn long-range interactions on sequential data, transformers continue to show state-of-the-art results on a wide variety of tasks. In contrast to CNNs, they contain no inductive bias that prioritizes local interactions. This makes them expressive, but also computationally infeasible for long sequences, such as high-resolution images. We demonstrate how combining the effectiveness of the inductive bias of CNNs with the expressivity of transformers enables them to model and thereby synthesize high-resolution images. We show how to (i) use CNNs to learn a contextrich vocabulary of image constituents, and in turn (ii) utilize transformers to efficiently model their composition within high-resolution images. Our approach is readily applied to conditional synthesis tasks, where both non-spatial information, such as object classes, and spatial information, such as segmentations, can control the generated image. In particular, we present the first results on semanticallyguided synthesis of megapixel images with transformers. Project page at https://git.io/JLlvY .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1,455
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- YOLOv10: Real-Time End-to-End Object DetectionAo Wang, Hui Chen, Lihao Liu, Kai Chen 等NeurIPS 2024 · 被引用 6,113 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
它引用的顶会 Paper17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine 等NeurIPS 2020 · 被引用 2,345 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 被引用 1,141 次
相关 Paper
- Geometry-Free View Synthesis: Transformers and no 3D PriorsRobin Rombach, Patrick Esser, Björn OmmerICCV 2021 · 被引用 115 次
- High-Fidelity Pluralistic Image Completion with TransformersZiyu Wan, Jingbo Zhang, Dongdong Chen, Jing LiaoICCV 2021 · 被引用 296 次
- Bridging Global Context Interactions for High-Fidelity Image CompletionChuanxia Zheng, Tat-Jen Cham, Jianfei Cai, Dinh Q. PhungCVPR 2022 · 被引用 114 次
- ReSTR: Convolution-free Referring Image Segmentation Using TransformersNamyup Kim, Dongwon Kim, Suha Kwak, Cuiling Lan 等CVPR 2022 · 被引用 149 次
- ASSET: autoregressive semantic scene editing with transformers at high resolutionsDifan Liu, Sandesh Shetty, Tobias Hinz, Matthew Fisher 等SIGGRAPH 2022 · 被引用 5 次
