Desigen: A Pipeline for Controllable Design Template Generation
Haohan Weng, Danqing Huang, Yu Qiao, Zheng Hu, Chin-Yew Lin, Tong Zhang, C. L. Philip Chen
Abstract
Templates serve as a good starting point to implement a design (e.g., banner, slide) but it takes great effort from designers to manually create. In this paper, we present Desigen, an automatic template creation pipeline which generates background images as well as harmonious layout elements over the background. Different from natural images, a background image should preserve enough non-salient space for the overlaying layout elements. To equip existing advanced diffusion-based models with stronger spatial control, we propose two simple but effective techniques to constrain the saliency distribution and reduce the attention weight in desired regions during the background generation process. Then conditioned on the background, we synthesize the layout with a Transformer-based autoregressive generator. To achieve a more harmonious composition, we propose an iterative inference strategy to adjust the synthesized background and layout in multiple rounds. We constructed a design dataset with more than 40k advertisement banners to verify our approach. Extensive experiments demonstrate that the proposed pipeline generates high-quality templates comparable to human designers. More than a single-page design, we further show an application of presentation generation that outputs a set of theme-consistent slides. The data and code are available at https://whaohan.github.io/desigen.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- NFIG: Multi-Scale Autoregressive Image Generation via Frequency OrderingZhihao Huang, Xi Qiu, Yukuo Ma, Yifu Zhou et al.NeurIPS 2025 · 20 citations
- CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image GenerationHui Zhang, Dexiang Hong, Yitong Wang, Jie Shao et al.ICCV 2025 · 7 citations
- BannerAgency: Advertising Banner Design with Multimodal LLM AgentsHeng Wang, Yotaro Shimose, Shingo TakamatsuEMNLP 2025 · 2 citations
- Uni-Layout: Integrating Human Feedback in Unified Layout Generation and EvaluationShuo Lu, Yanyin Chen, Wei Feng, Jiahao Fan et al.ACM MM 2025 · 1 citation
- Masked Region Transformer for Layered Image Generation and Editing at ScaleZhicong Tang, Jingye Chen, Zhao Zhang, Mohan Zhou et al.CVPR 2026
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Constrained Graphic Layout Generation via Latent OptimizationKotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, Kota YamaguchiACM MM 2021 · 80 citations
- LayoutDM: Transformer-based Diffusion Model for Layout GenerationShang Chai, Liansheng Zhuang, Fengying YanCVPR 2023
- Geometry Aligned Variational Transformer for Image-conditioned Layout GenerationYunning Cao, Ye Ma, Min Zhou, Chuanbin Liu et al.ACM MM 2022 · 37 citations
- CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic DesignHui Zhang, Dexiang Hong, Maoke Yang, Yutao Cheng et al.ICLR 2026 · 40 citations
- DreamFuse: Adaptive Image Fusion with Diffusion TransformerJunjia Huang, Pengxiang Yan, Jiyang Liu, Jie Wu et al.ICCV 2025 · 3 citations
