Desigen: A Pipeline for Controllable Design Template Generation
Haohan Weng, Danqing Huang, Yu Qiao, Zheng Hu, Chin-Yew Lin, Tong Zhang, C. L. Philip Chen
摘要
Templates serve as a good starting point to implement a design (e.g., banner, slide) but it takes great effort from designers to manually create. In this paper, we present Desigen, an automatic template creation pipeline which generates background images as well as harmonious layout elements over the background. Different from natural images, a background image should preserve enough non-salient space for the overlaying layout elements. To equip existing advanced diffusion-based models with stronger spatial control, we propose two simple but effective techniques to constrain the saliency distribution and reduce the attention weight in desired regions during the background generation process. Then conditioned on the background, we synthesize the layout with a Transformer-based autoregressive generator. To achieve a more harmonious composition, we propose an iterative inference strategy to adjust the synthesized background and layout in multiple rounds. We constructed a design dataset with more than 40k advertisement banners to verify our approach. Extensive experiments demonstrate that the proposed pipeline generates high-quality templates comparable to human designers. More than a single-page design, we further show an application of presentation generation that outputs a set of theme-consistent slides. The data and code are available at https://whaohan.github.io/desigen.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- NFIG: Multi-Scale Autoregressive Image Generation via Frequency OrderingZhihao Huang, Xi Qiu, Yukuo Ma, Yifu Zhou 等NeurIPS 2025 · 被引用 20 次
- CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image GenerationHui Zhang, Dexiang Hong, Yitong Wang, Jie Shao 等ICCV 2025 · 被引用 7 次
- BannerAgency: Advertising Banner Design with Multimodal LLM AgentsHeng Wang, Yotaro Shimose, Shingo TakamatsuEMNLP 2025 · 被引用 2 次
- Uni-Layout: Integrating Human Feedback in Unified Layout Generation and EvaluationShuo Lu, Yanyin Chen, Wei Feng, Jiahao Fan 等ACM MM 2025 · 被引用 1 次
- Masked Region Transformer for Layered Image Generation and Editing at ScaleZhicong Tang, Jingye Chen, Zhao Zhang, Mohan Zhou 等CVPR 2026
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Constrained Graphic Layout Generation via Latent OptimizationKotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, Kota YamaguchiACM MM 2021 · 被引用 80 次
- LayoutDM: Transformer-based Diffusion Model for Layout GenerationShang Chai, Liansheng Zhuang, Fengying YanCVPR 2023
- Geometry Aligned Variational Transformer for Image-conditioned Layout GenerationYunning Cao, Ye Ma, Min Zhou, Chuanbin Liu 等ACM MM 2022 · 被引用 37 次
- CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic DesignHui Zhang, Dexiang Hong, Maoke Yang, Yutao Cheng 等ICLR 2026 · 被引用 40 次
- DreamFuse: Adaptive Image Fusion with Diffusion TransformerJunjia Huang, Pengxiang Yan, Jiyang Liu, Jie Wu 等ICCV 2025 · 被引用 3 次
