From Elements to Design: A Layered Approach for Automatic Graphic Design Composition
Jiawei Lin, Shizhao Sun, Danqing Huang, Ting Liu, Ji Li, Jiang Bian
Abstract
In this work, we investigate automatic design composition from multimodal graphic elements. Although recent studies have developed various generative models for graphic design, they usually face the following limitations: they only focus on certain subtasks and are far from achieving the design composition task; they do not consider the hierarchical information of graphic designs during the generation process. To tackle these issues, we introduce the layered design principle into Large Multimodal Models (LMMs) and propose a novel approach, called LaDeCo, to accomplish this challenging task. Specifically, LaDeCo first performs layer planning for a given element set, dividing the input elements into different semantic layers according to their contents. Based on the planning results, it subsequently predicts element attributes that control the design composition in a layer-wise manner, and includes the rendered image of previously generated layers into the context. With this insightful design, LaDeCo decomposes the difficult task into smaller manageable steps, making the genera-tion process smoother and clearer. The experimental results demonstrate the effectiveness of LaDeCo in design composition. Furthermore, we show that LaDeCo enables some interesting applications in graphic design, such as resolution adjustment, design decoration, design variation, etc. In addition, it even outperforms the specialized models in some design subtasks without any task-specific training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- LayerD: Decomposing Raster Graphic Designs into LayersTomoyuki Suzuki, Kang-Jun Liu, Naoto Inoue, Kota YamaguchiICCV 2025 · 3 citations
- BannerAgency: Advertising Banner Design with Multimodal LLM AgentsHeng Wang, Yotaro Shimose, Shingo TakamatsuEMNLP 2025 · 2 citations
- Seeing is Improving: Visual Feedback for Iterative Text Layout RefinementJunrong Guo, Shancheng Fang, Yadong Qu, Hongtao XieCVPR 2026 · 2 citations
- PSDesigner: Automated Graphic Design with a Human-Like Creative WorkflowXincheng Shuai, Song Tang, Yutong Huang, Henghui Ding et al.CVPR 2026 · 1 citation
- AnyDoc: Enhancing Document Generation via Large-Scale HTML/CSS Data Synthesis and Height-Aware Reinforcement OptimizationJiawei Lin, Wanrong Zhu, Vlad I Morariu, Christopher TensmeyerCVPR 2026 · 1 citation
Builds on20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
- CanvasVAE: Learning to Generate Vector Graphic DocumentsKota YamaguchiICCV 2021 · 103 citations
Related papers
- Graphic Design with Large Multimodal ModelYutao Cheng, Zhao Zhang, Maoke Yang, Hui Nie et al.AAAI 2025 · 3 citations
- Decomposition of Graphic Design with Unified Multimodal ModelHui Nie, Zhao Zhang, Yutao Cheng, Maoke Yang et al.ICML 2025
- Rethinking Layered Graphic Design Generation with a Top-Down ApproachJingye Chen, Zhaowen Wang, Nanxuan Zhao, Li Zhang et al.ICCV 2025 · 4 citations
- Igd: Instructional Graphic Design With Multimodal Layer GeneratioYadong Qu, Hongtao Xie, Yongdong Zhang, Shancheng Fang et al.ICCV 2025 · 1 citation
- Multimodal Markup Document Models for Graphic Design CompletionKotaro Kikuchi, Ukyo Honda, Naoto Inoue, Mayu Otani et al.ACM MM 2025 · 1 citation
