SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion Models
Jaerin Lee, Daniel Sungho Jung, Kanggeon Lee, Kyoung Mu Lee
Abstract
We introduce SemanticDraw, a new paradigm of interactive content creation where high-quality images are generated in near real-time from given multiple hand-drawn regions, each encoding prescribed semantic meaning. In order to maximize the productivity of content creators and to fully realize their artistic imagination, it requires both quick interactive interfaces and fine-grained regional controls in their tools. Despite astonishing generation quality from recent diffusion models, we find that existing approaches for regional controllability are very slow (52 seconds for 512 × 512 image) while not compatible with acceleration methods such as LCM, blocking their huge potential in interactive content creation. From this observation, we build our solution for interactive content creation in two steps: (1) we establish compatibility between region-based controls and acceleration techniques for diffusion models, maintaining high fidelity of multi-prompt image generation with ×10 reduced number of inference steps, (2) we increase the generation throughput with our new multi-prompt stream batch pipeline, enabling low-latency generation from multiple, region-based text prompts on a single RTX 2080 Ti GPU. Our proposed framework is generalizable to any existing diffusion models and acceleration schedulers, allowing sub-second (0.64 seconds) image content creation application upon well-established image diffusion models. The code is https://github.com/ironjr/semantic-draw.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Learning Dense Hand Contact Estimation from Imbalanced DataDaniel Sungho Jung, Kyoung Mu LeeNeurIPS 2025 · 14 citations
- GeoRK2: Geometry-Guided Runge–Kutta Integration for Diffusion Transformer AccelerationChaoqun Sun, Zongjing Fu, Powei Chang, Jinpeng Zhang et al.CVPR 2026
- PHAC: Promptable Human Amodal CompletionSeung Young Noh, Ju Yong ChangCVPR 2026
Builds on35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- StreamDiffusion: A Pipeline-Level Solution for Real-Time Interactive GenerationAkio Kodaira, Chenfeng Xu, Toshiki Hazama, Takanori Yoshimoto et al.ICCV 2025 · 10 citations
- Generating compositional scenes via Text-to-image RGBA Instance GenerationAlessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang et al.NeurIPS 2024 · 13 citations
- Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic PromptsXinhua Cheng, Tianyu Yang, Jianan Wang, Yu Li et al.ICLR 2024 · 58 citations
- DialogDraw: Image Generation and Editing System Based on Multi-Turn DialogueShichao Ma, Xinfeng Zhang, Zeng Zhao, Bai Liu et al.AAAI 2025 · 3 citations
- Diffusion Texture PaintingAnita Hu, Nishkrit Desai, Hassan Abu Alhaija, Seung Wook Kim et al.SIGGRAPH 2024 · 16 citations
