Hierarchical Image Generation via Transformer-Based Sequential Patch Selection
Xiaogang Xu, Ning Xu
Abstract
To synthesize images with preferred objects and interactions, a controllable way is to generate the image from a scene graph and a large pool of object crops, where the spatial arrangements of the objects in the image are defined by the scene graph while their appearances are determined by the retrieved crops from the pool. In this paper, we propose a novel framework with such a semi-parametric generation strategy. First, to encourage the retrieval of mutually compatible crops, we design a sequential selection strategy where the crop selection for each object is determined by the contents and locations of all object crops that have been chosen previously. Such process is implemented via a transformer trained with contrastive losses. Second, to generate the final image, our hierarchical generation strategy leverages hierarchical gated convolutions which are employed to synthesize areas not covered by any image crops, and a patch guided spatially adaptive normalization module which is proposed to guarantee the final generated images complying with the crop appearance and the scene graph. Evaluated on the challenging Visual Genome and COCO-Stuff dataset, our experimental results demonstrate the superiority of our proposed method over existing state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75359e73-226f-4395-8f1d-4700fec054e2Cited by top-tier papers1
Ask how each one uses itBuilds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 190 citations
- Self-Supervised Sketch-to-Image SynthesisBingchen Liu, Yizhe Zhu, Kunpeng Song, Ahmed ElgammalAAAI 2021 · 46 citations
Related papers
- R3CD: Scene Graph to Image Generation with Relation-Aware Compositional Contrastive Control DiffusionJinxiu Liu, Qi LiuAAAI 2024 · 21 citations
- LayoutTransformer: Scene Layout Generation With Conceptual and Spatial DiversityCheng-Fu Yang, Wan-Cyuan Fan, Fu-En Yang, Yu-Chiang Frank WangCVPR 2021
- Scene Graph Expansion for Semantics-Guided Image OutpaintingChiao-An Yang, Cheng-Yo Tan, Wan-Cyuan Fan, Cheng-Fu Yang et al.CVPR 2022 · 18 citations
- Fully Functional Image Manipulation Using Scene Graphs in A Bounding-Box Free WaySitong Su, Lianli Gao, Junchen Zhu, Jie Shao et al.ACM MM 2021 · 16 citations
- Object-Centric Image Generation from LayoutsTristan Sylvain, Pengchuan Zhang, Yoshua Bengio, R. Devon Hjelm et al.AAAI 2021 · 107 citations
