Text-to-Image Synthesis based on Object-Guided Joint-Decoding Transformer
Fuxiang Wu, Liu Liu, Fusheng Hao, Fengxiang He, Jun Cheng
Abstract
Object-guided text-to-image synthesis aims to generate images from natural language descriptions built by two-step frameworks, i.e., the model generates the layout and then synthesizes images from the layout and captions. However, such frameworks have two issues: 1) complex structure, since generating language-related layout is not a trivial task; 2) error propagation, because the inappropriate layout will mislead the image synthesis and is hard to be revised. In this paper, we propose an object-guided joint-decoding module to simultaneously generate the image and the corresponding layout. Specially, we present the joint-decoding transformer to model the joint probability on images tokens and the corresponding layouts tokens, where layout tokens provide additional observed data to model the complex scene better. Then, we describe a novel Layout-Vqgan for layout encoding and decoding to provide more information about the complex scene. After that, we present the detail-enhanced module to enrich the language-related details based on two facts: 1) visual details could be omitted in the compression of VQGANs; 2) the joint-decoding transformer would not have sufficient generating capacity. The experiments show that our approach is competitive with previous object-centered models and can generate diverse and high-quality objects under the given layouts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3575d7e-86a2-459a-bbb1-6d586e0f658dCited by top-tier papers2
- Reject Decoding via Language-Vision Models for Text-to-Image SynthesisFuxiang Wu, Liu Liu, Fusheng Hao, Fengxiang He et al.AAAI 2023 · 2 citations
- Toward Verifiable and Reproducible Human Evaluation for Text-to-Image GenerationMayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami et al.CVPR 2023
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng et al.NeurIPS 2021 · 1,026 citations
- Image Synthesis From Reconfigurable Layout and StyleWei Sun, Tianfu WuICCV 2019 · 160 citations
Related papers
- PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language ModelsRunze He, Bo Cheng, Yuhang Ma, Qingxiang Jia et al.ICCV 2025 · 1 citation
- Rethinking the Objectives of Vector-Quantized Tokenizers for Image SynthesisYuchao Gu, Xintao Wang, Yixiao Ge, Ying Shan et al.CVPR 2024
- Compositional Transformers for Scene GenerationDrew A. Hudson, Larry ZitnickNeurIPS 2021 · 36 citations
- Background Layout Generation and Object Knowledge Transfer for Text-to-Image GenerationZhuowei Chen, Zhendong Mao, Shancheng Fang, Bo HuACM MM 2022 · 6 citations
- NÜWA-LIP: Language-guided Image Inpainting with Defect-free VQGANMinheng Ni, Xiaoming Li, Wangmeng ZuoCVPR 2023
