Compositional Transformers for Scene Generation
Drew A. Hudson, Larry Zitnick
Abstract
We introduce the GANformer2 model, an iterative object-oriented transformer, explored for the task of generative modeling. The network incorporates strong and explicit structural priors, to reflect the compositional nature of visual scenes, and synthesizes images through a sequential process. It operates in two stages: a fast and lightweight planning phase, where we draft a high-level scene layout, followed by an attention-based execution phase, where the layout is being refined, evolving into a rich and detailed picture. Our model moves away from conventional black-box GAN architectures that feature a flat and monolithic latent space towards a transparent design that encourages efficiency, controllability and interpretability. We demonstrate GANformer2's strengths and qualities through a careful evaluation over a range of datasets, from multi-object CLEVR scenes to the challenging COCO images, showing it successfully achieves state-of-the-art performance in terms of visual quality, diversity and consistency. Further experiments demonstrate the model's disentanglement and provide a deeper insight into its generative process, as it proceeds step-by-step from a rough initial sketch, to a detailed layout that accounts for objects' depths and dependencies, and up to the final high-resolution depiction of vibrant and intricate real-world scenes. See https://github.com/ dorarad/gansformer for model implementation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32f110fc-24c6-4625-9327-b5b1a97f1939Cited by top-tier papers9
- Class-Aware Adversarial Transformers for Medical Image SegmentationChenyu You, Ruihan Zhao, Fenglin Liu, Siyuan Dong et al.NeurIPS 2022 · 137 citations
- SemanticStyleGAN: Learning Compositional Generative Priors for Controllable Image Synthesis and EditingYichun Shi, Xiao Yang, Yangyue Wan, Xiaohui ShenCVPR 2022 · 88 citations
- CC3D: Layout-Conditioned Generation of Compositional 3D ScenesSherwin Bahmani, Jeong Joon Park, Despoina Paschalidou, Xingguang Yan et al.ICCV 2023 · 66 citations
- LinkGAN: Linking GAN Latents to Pixels for Controllable Image SynthesisJiapeng Zhu, Ceyuan Yang, Yujun Shen, Zifan Shi et al.ICCV 2023 · 29 citations
- Semantic 3D-Aware Portrait Synthesis and Manipulation Based on Compositional Neural Radiance FieldTianxiang Ma, Bingchuan Li, Qian He, Jing Dong et al.AAAI 2023 · 13 citations
Builds on19
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 459 citations
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 421 citations
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 334 citations
Related papers
- Generative Adversarial TransformersDrew A. Hudson, Larry ZitnickICML 2021 · 213 citations
- Text-to-Image Synthesis based on Object-Guided Joint-Decoding TransformerFuxiang Wu, Liu Liu, Fusheng Hao, Fengxiang He et al.CVPR 2022 · 13 citations
- Styleformer: Transformer based Generative Adversarial Networks with Style VectorJeeseung Park, Younggeun KimCVPR 2022 · 49 citations
- LayoutTransformer: Scene Layout Generation With Conceptual and Spatial DiversityCheng-Fu Yang, Wan-Cyuan Fan, Fu-En Yang, Yu-Chiang Frank WangCVPR 2021
- Image Synthesis From Reconfigurable Layout and StyleWei Sun, Tianfu WuICCV 2019 · 160 citations
