Igd: Instructional Graphic Design With Multimodal Layer Generatio
Yadong Qu, Hongtao Xie, Yongdong Zhang, Shancheng Fang, Yuxin Wang, Xiaorui Wang, Zhineng Chen
摘要
Graphic design visually conveys information and data by creating and combining text, images and graphics. Twostage methods that rely primarily on layout generation lack creativity and intelligence, making graphic design still labor-intensive. Existing diffusion-based methods generate non-editable graphic design files at image level with poor legibility in visual text rendering, which prevents them from achieving satisfactory and practical automated graphic design. In this paper, we propose Instructional Graphic Designer (IGD) to swiftly generate multimodal layers with editable flexibility with only natural language instructions. IGD adopts a new paradigm that leverages parametric rendering and image asset generation. First, we develop a design platform and establish a standardized format for multiscenario design files, thus laying the foundation for scaling up data. Second, IGD utilizes the multimodal understanding and reasoning capabilities of MLLM to accomplish attribute prediction, sequencing and layout of layers. It also employs a diffusion model to generate image content for assets. By enabling end-to-end training, IGD architecturally supports scalability and extensibility in complex graphic design tasks. The superior experimental results demonstrate that IGD offers a new solution for graphic design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- PSDesigner: Automated Graphic Design with a Human-Like Creative WorkflowXincheng Shuai, Song Tang, Yutong Huang, Henghui Ding 等CVPR 2026 · 被引用 1 次
- Multimodal Markup Document Models for Graphic Design CompletionKotaro Kikuchi, Ukyo Honda, Naoto Inoue, Mayu Otani 等ACM MM 2025 · 被引用 1 次
- Graphic Design with Large Multimodal ModelYutao Cheng, Zhao Zhang, Maoke Yang, Hui Nie 等AAAI 2025 · 被引用 3 次
- Rethinking Layered Graphic Design Generation with a Top-Down ApproachJingye Chen, Zhaowen Wang, Nanxuan Zhao, Li Zhang 等ICCV 2025 · 被引用 4 次
- CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic DesignHui Zhang, Dexiang Hong, Maoke Yang, Yutao Cheng 等ICLR 2026 · 被引用 40 次
