Illustration Layout Generation for Slide Enhancement with Pixel-based Diffusion Model
Zhaoyun Jiang, Jiaqi Guo, Shakie Liu, Chao Han, Ting Liu, Jian-Guang Lou, Dongmei Zhang
Abstract
Embellishing slides with illustrations is a well-established practice for improving engagement and storytelling. However, this process is challenging, requiring careful consideration of both visual appearance and semantics of illustrations while ensuring they complement rather than overwhelm the slide content. In this paper, we take a pioneering step toward automating this process by introducing the task of Illustration Layout Generation: given a slide and a set of illustrations, automatically determining their optimal sizes and positions to enrich the slide. Existing layout generation approaches struggle with this task as they rely on large-scale layout datasets for training and have limited support for multiple visual inputs. To address these challenges, we propose SlideILG, a method that iteratively optimizes illustration placement using a diffusion-based text-to-image prior. We introduce three key techniques to enhance efficiency and quality: (1) leveraging cross-attention maps from the text-to-image model to initialize illustration placement; (2) employing an over-parameterization strategy to stabilize optimization; and (3) fine-tuning the text-to-image model on high-quality slide thumbnails for more precise guidance. To evaluate SlideILG, we construct IllustrationBench, a benchmark comprising 128 real-world slides, each paired with a set of illustrations for embellishment. Quantitative, qualitative and human-study results demonstrate the effectiveness of our approach. Furthermore, we showcase a real-world application scenario to highlight the significance and practical utility of this task and our method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 71992d93-66be-4de1-b0ff-39f520c94a04Related papers
- Desigen: A Pipeline for Controllable Design Template GenerationHaohan Weng, Danqing Huang, Yu Qiao, Zheng Hu et al.CVPR 2024 · 10 citations
- AutoFigure: Generating and Refining Publication-Ready Scientific IllustrationsMinjun Zhu, Zhen Lin, Yixuan Weng, Panzhong Lu et al.ICLR 2026 · 28 citations
- AI Illustrator: Translating Raw Descriptions into Images by Prompt-based Cross-Modal GenerationYiyang Ma, Huan Yang, Bei Liu, Jianlong Fu et al.ACM MM 2022 · 9 citations
- Grounded Text-to-Image Synthesis with Attention RefocusingQuynh Phung, Songwei Ge, Jia-Bin HuangCVPR 2024 · 59 citations
- SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from DesignWenxin Tang, Jingyu Xiao, Wenxuan Jiang, Xi Xiao et al.EMNLP 2025 · 2 citations
