From Elements to Design: A Layered Approach for Automatic Graphic Design Composition
Jiawei Lin, Shizhao Sun, Danqing Huang, Ting Liu, Ji Li, Jiang Bian
摘要
In this work, we investigate automatic design composition from multimodal graphic elements. Although recent studies have developed various generative models for graphic design, they usually face the following limitations: they only focus on certain subtasks and are far from achieving the design composition task; they do not consider the hierarchical information of graphic designs during the generation process. To tackle these issues, we introduce the layered design principle into Large Multimodal Models (LMMs) and propose a novel approach, called LaDeCo, to accomplish this challenging task. Specifically, LaDeCo first performs layer planning for a given element set, dividing the input elements into different semantic layers according to their contents. Based on the planning results, it subsequently predicts element attributes that control the design composition in a layer-wise manner, and includes the rendered image of previously generated layers into the context. With this insightful design, LaDeCo decomposes the difficult task into smaller manageable steps, making the genera-tion process smoother and clearer. The experimental results demonstrate the effectiveness of LaDeCo in design composition. Furthermore, we show that LaDeCo enables some interesting applications in graphic design, such as resolution adjustment, design decoration, design variation, etc. In addition, it even outperforms the specialized models in some design subtasks without any task-specific training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- LayerD: Decomposing Raster Graphic Designs into LayersTomoyuki Suzuki, Kang-Jun Liu, Naoto Inoue, Kota YamaguchiICCV 2025 · 被引用 3 次
- BannerAgency: Advertising Banner Design with Multimodal LLM AgentsHeng Wang, Yotaro Shimose, Shingo TakamatsuEMNLP 2025 · 被引用 2 次
- Seeing is Improving: Visual Feedback for Iterative Text Layout RefinementJunrong Guo, Shancheng Fang, Yadong Qu, Hongtao XieCVPR 2026 · 被引用 2 次
- PSDesigner: Automated Graphic Design with a Human-Like Creative WorkflowXincheng Shuai, Song Tang, Yutong Huang, Henghui Ding 等CVPR 2026 · 被引用 1 次
- AnyDoc: Enhancing Document Generation via Large-Scale HTML/CSS Data Synthesis and Height-Aware Reinforcement OptimizationJiawei Lin, Wanrong Zhu, Vlad I Morariu, Christopher TensmeyerCVPR 2026 · 被引用 1 次
它引用的顶会 Paper20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- CanvasVAE: Learning to Generate Vector Graphic DocumentsKota YamaguchiICCV 2021 · 被引用 103 次
相关 Paper
- Graphic Design with Large Multimodal ModelYutao Cheng, Zhao Zhang, Maoke Yang, Hui Nie 等AAAI 2025 · 被引用 3 次
- Decomposition of Graphic Design with Unified Multimodal ModelHui Nie, Zhao Zhang, Yutao Cheng, Maoke Yang 等ICML 2025
- Rethinking Layered Graphic Design Generation with a Top-Down ApproachJingye Chen, Zhaowen Wang, Nanxuan Zhao, Li Zhang 等ICCV 2025 · 被引用 4 次
- Igd: Instructional Graphic Design With Multimodal Layer GeneratioYadong Qu, Hongtao Xie, Yongdong Zhang, Shancheng Fang 等ICCV 2025 · 被引用 1 次
- Multimodal Markup Document Models for Graphic Design CompletionKotaro Kikuchi, Ukyo Honda, Naoto Inoue, Mayu Otani 等ACM MM 2025 · 被引用 1 次
