Geometry Aligned Variational Transformer for Image-conditioned Layout Generation
Yunning Cao, Ye Ma, Min Zhou, Chuanbin Liu, Hongtao Xie, Tiezheng Ge, Yuning Jiang
Abstract
Layout generation is a novel task in computer vision, which combines the challenges in both object localization and aesthetic appraisal, widely used in advertisements, posters, and slides design. An accurate and pleasant layout should consider both the intradomain relationship within layout elements and the inter-domain relationship between layout elements and the image. However, most previous methods simply focus on image-content-agnostic layout generation, without leveraging the complex visual information from the image. To this end, we explore a novel paradigm entitled image-conditioned layout generation, which aims to add text overlays to an image in a semantically coherent manner. Specifically, we propose an Image-Conditioned Variational Transformer (ICVT) that autoregressively generates various layouts in an image. First, self-attention mechanism is adopted to model the contextual relationship within layout elements, while cross-attention mechanism is used to fuse the visual information of conditional images. Subsequently, we take them as building blocks of conditional variational autoencoder (CVAE), which demonstrates appealing diversity. Second, in order to alleviate the gap between layout elements domain and visual domain, we design a Geometry Alignment module, in which the geometric information of the image is aligned with the layout representation. In addition, we construct a large-scale advertisement poster layout designing dataset with delicate layout and saliency map annotations. Experimental results show that our model can adaptively generate layouts in the non-intrusive area of the image, resulting in a harmonious layout design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66f01013-bfbf-45ba-bdd9-c42df9720ab6Cited by top-tier papers18
- PLay: Parametrically Conditioned Layout Generation using Latent DiffusionChin-Yi Cheng, Forrest Huang, Gang Li, Yang LiICML 2023 · 45 citations
- AutoPoster: A Highly Automatic and Content-aware Design System for Advertising Poster GenerationJinpeng Lin, Min Zhou, Ye Ma, Yifan Gao et al.ACM MM 2023 · 24 citations
- Tell2Design: A Dataset for Language-Guided Floor Plan GenerationSicong Leng, Yang Zhou, Mohammed Haroon Dupty, Wee Sun Lee et al.ACL 2023 · 20 citations
- Desigen: A Pipeline for Controllable Design Template GenerationHaohan Weng, Danqing Huang, Yu Qiao, Zheng Hu et al.CVPR 2024 · 10 citations
- TextPainter: Multimodal Text Image Generation with Visual-harmony and Text-comprehension for Poster DesignYifan Gao, Jinpeng Lin, Min Zhou, Chuanbin Liu et al.ACM MM 2023 · 6 citations
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- LayoutVAE: Stochastic Scene Layout Generation From a Label SetAkash Abdu Jyothi, Thibaut Durand, Jiawei He, Leonid Sigal et al.ICCV 2019 · 194 citations
- LayoutTransformer: Layout Generation and Completion with Self-attentionKamal Gupta, Justin Lazarow, Alessandro Achille, Larry Davis et al.ICCV 2021 · 184 citations
Related papers
- Variational Transformer Networks for Layout GenerationDiego Martín Arroyo, Janis Postels, Federico TombariCVPR 2021
- Constrained Graphic Layout Generation via Latent OptimizationKotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, Kota YamaguchiACM MM 2021 · 80 citations
- PosterLayout: A New Benchmark and Approach for Content-Aware Visual-Textual Presentation LayoutHsiaoYuan Hsu, Xiangteng He, Yuxin Peng, Hao Kong et al.CVPR 2023
- Unsupervised Domain Adaption with Pixel-Level Discriminator for Image-Aware Layout GenerationChenchen Xu, Min Zhou, Tiezheng Ge, Yuning Jiang et al.CVPR 2023
- LayoutTransformer: Scene Layout Generation With Conceptual and Spatial DiversityCheng-Fu Yang, Wan-Cyuan Fan, Fu-En Yang, Yu-Chiang Frank WangCVPR 2021
