CanvasEmb: Learning Layout Representation with Large-scale Pre-training for Graphic Design
Yuxi Xie, Danqing Huang, Jinpeng Wang, Chin-Yew Lin
Abstract
Layout representation, which models visual elements and their inter-relations in a canvas, plays a crucial role in graphic design intelligence. With a large variety of layout designs and the unique characteristic of layouts that visual elements are defined as a list of categorical (e.g., type) and numerical (e.g., position and size) properties, it is challenging to learn general and compact representations with limited data. Inspired by the recent success of self-supervised pre-training techniques in various natural language processing tasks, in this paper, we propose CanvasEmb (Canvas Embedding), which pre-trains deep representations from unlabeled graphic designs by jointly conditioning on all the context elements in a canvas, with a multi-dimensional feature encoder and a multi-task learning objective. The pre-trained CanvasEmb model can be fine-tuned with just one additional output layer and with a small size of training data to create models for a wide range of downstream tasks. We verify our approach with presentation slides data. We construct a large-scale dataset with more than one million slides and propose two layout understanding tasks with human-labeled sets, namely element role labeling and image captioning. Evaluation results on these two tasks show that our model with fine-tuning achieves state-of-the-art performance. Furthermore, we conduct a deep analysis aiming to understand the modeling mechanism of CanvasEmb and demonstrate its great potential with two extended applications: layout auto completion and layout retrieval.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 96771bf7-715f-496a-8d5f-2bddc4af87fcCited by top-tier papers2
- Demystifying Tacit Knowledge in Graphic Design: Characteristics, Instances, Approaches, and GuidelinesKihoon Son, DaEun Choi, Tae Soo Kim, Juho KimCHI 2024 · 14 citations
- Spot the Error: Non-autoregressive Graphic Layout Generation with Wireframe LocatorJieru Lin, Danqing Huang, Tiejun Zhao, Dechen Zhan et al.AAAI 2024 · 5 citations
Related papers
- CanvasVAE: Learning to Generate Vector Graphic DocumentsKota YamaguchiICCV 2021 · 103 citations
- Screen2Vec: Semantic Embedding of GUI Screens and GUI ComponentsToby Jia-Jun Li, Lindsay Popowski, Tom M. Mitchell, Brad A. MyersCHI 2021 · 72 citations
- UniDoc: Unified Pretraining Framework for Document UnderstandingJiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao et al.NeurIPS 2021 · 118 citations
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
