Compositional Transformers for Scene Generation
Drew A. Hudson, Larry Zitnick
摘要
We introduce the GANformer2 model, an iterative object-oriented transformer, explored for the task of generative modeling. The network incorporates strong and explicit structural priors, to reflect the compositional nature of visual scenes, and synthesizes images through a sequential process. It operates in two stages: a fast and lightweight planning phase, where we draft a high-level scene layout, followed by an attention-based execution phase, where the layout is being refined, evolving into a rich and detailed picture. Our model moves away from conventional black-box GAN architectures that feature a flat and monolithic latent space towards a transparent design that encourages efficiency, controllability and interpretability. We demonstrate GANformer2's strengths and qualities through a careful evaluation over a range of datasets, from multi-object CLEVR scenes to the challenging COCO images, showing it successfully achieves state-of-the-art performance in terms of visual quality, diversity and consistency. Further experiments demonstrate the model's disentanglement and provide a deeper insight into its generative process, as it proceeds step-by-step from a rough initial sketch, to a detailed layout that accounts for objects' depths and dependencies, and up to the final high-resolution depiction of vibrant and intricate real-world scenes. See https://github.com/ dorarad/gansformer for model implementation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Class-Aware Adversarial Transformers for Medical Image SegmentationChenyu You, Ruihan Zhao, Fenglin Liu, Siyuan Dong 等NeurIPS 2022 · 被引用 137 次
- SemanticStyleGAN: Learning Compositional Generative Priors for Controllable Image Synthesis and EditingYichun Shi, Xiao Yang, Yangyue Wan, Xiaohui ShenCVPR 2022 · 被引用 88 次
- CC3D: Layout-Conditioned Generation of Compositional 3D ScenesSherwin Bahmani, Jeong Joon Park, Despoina Paschalidou, Xingguang Yan 等ICCV 2023 · 被引用 66 次
- LinkGAN: Linking GAN Latents to Pixels for Controllable Image SynthesisJiapeng Zhu, Ceyuan Yang, Yujun Shen, Zifan Shi 等ICCV 2023 · 被引用 29 次
- Semantic 3D-Aware Portrait Synthesis and Manipulation Based on Compositional Neural Radiance FieldTianxiang Ma, Bingchuan Li, Qian He, Jing Dong 等AAAI 2023 · 被引用 13 次
它引用的顶会 Paper19
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 被引用 1,049 次
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 被引用 459 次
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 被引用 421 次
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 被引用 334 次
相关 Paper
- Generative Adversarial TransformersDrew A. Hudson, Larry ZitnickICML 2021 · 被引用 213 次
- Text-to-Image Synthesis based on Object-Guided Joint-Decoding TransformerFuxiang Wu, Liu Liu, Fusheng Hao, Fengxiang He 等CVPR 2022 · 被引用 13 次
- Styleformer: Transformer based Generative Adversarial Networks with Style VectorJeeseung Park, Younggeun KimCVPR 2022 · 被引用 49 次
- LayoutTransformer: Scene Layout Generation With Conceptual and Spatial DiversityCheng-Fu Yang, Wan-Cyuan Fan, Fu-En Yang, Yu-Chiang Frank WangCVPR 2021
- Image Synthesis From Reconfigurable Layout and StyleWei Sun, Tianfu WuICCV 2019 · 被引用 160 次
