Igd: Instructional Graphic Design With Multimodal Layer Generatio
Yadong Qu, Hongtao Xie, Yongdong Zhang, Shancheng Fang, Yuxin Wang, Xiaorui Wang, Zhineng Chen
Abstract
Graphic design visually conveys information and data by creating and combining text, images and graphics. Twostage methods that rely primarily on layout generation lack creativity and intelligence, making graphic design still labor-intensive. Existing diffusion-based methods generate non-editable graphic design files at image level with poor legibility in visual text rendering, which prevents them from achieving satisfactory and practical automated graphic design. In this paper, we propose Instructional Graphic Designer (IGD) to swiftly generate multimodal layers with editable flexibility with only natural language instructions. IGD adopts a new paradigm that leverages parametric rendering and image asset generation. First, we develop a design platform and establish a standardized format for multiscenario design files, thus laying the foundation for scaling up data. Second, IGD utilizes the multimodal understanding and reasoning capabilities of MLLM to accomplish attribute prediction, sequencing and layout of layers. It also employs a diffusion model to generate image content for assets. By enabling end-to-end training, IGD architecturally supports scalability and extensibility in complex graphic design tasks. The superior experimental results demonstrate that IGD offers a new solution for graphic design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b824a98e-0198-4de9-8754-4fdb41f7f54cCited by top-tier papers1
Ask how each one uses itBuilds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- PSDesigner: Automated Graphic Design with a Human-Like Creative WorkflowXincheng Shuai, Song Tang, Yutong Huang, Henghui Ding et al.CVPR 2026 · 1 citation
- Multimodal Markup Document Models for Graphic Design CompletionKotaro Kikuchi, Ukyo Honda, Naoto Inoue, Mayu Otani et al.ACM MM 2025 · 1 citation
- Graphic Design with Large Multimodal ModelYutao Cheng, Zhao Zhang, Maoke Yang, Hui Nie et al.AAAI 2025 · 3 citations
- Rethinking Layered Graphic Design Generation with a Top-Down ApproachJingye Chen, Zhaowen Wang, Nanxuan Zhao, Li Zhang et al.ICCV 2025 · 4 citations
- CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic DesignHui Zhang, Dexiang Hong, Maoke Yang, Yutao Cheng et al.ICLR 2026 · 40 citations
