SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion Models
Jaerin Lee, Daniel Sungho Jung, Kanggeon Lee, Kyoung Mu Lee
摘要
We introduce SemanticDraw, a new paradigm of interactive content creation where high-quality images are generated in near real-time from given multiple hand-drawn regions, each encoding prescribed semantic meaning. In order to maximize the productivity of content creators and to fully realize their artistic imagination, it requires both quick interactive interfaces and fine-grained regional controls in their tools. Despite astonishing generation quality from recent diffusion models, we find that existing approaches for regional controllability are very slow (52 seconds for 512 × 512 image) while not compatible with acceleration methods such as LCM, blocking their huge potential in interactive content creation. From this observation, we build our solution for interactive content creation in two steps: (1) we establish compatibility between region-based controls and acceleration techniques for diffusion models, maintaining high fidelity of multi-prompt image generation with ×10 reduced number of inference steps, (2) we increase the generation throughput with our new multi-prompt stream batch pipeline, enabling low-latency generation from multiple, region-based text prompts on a single RTX 2080 Ti GPU. Our proposed framework is generalizable to any existing diffusion models and acceleration schedulers, allowing sub-second (0.64 seconds) image content creation application upon well-established image diffusion models. The code is https://github.com/ironjr/semantic-draw.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning Dense Hand Contact Estimation from Imbalanced DataDaniel Sungho Jung, Kyoung Mu LeeNeurIPS 2025 · 被引用 14 次
- GeoRK2: Geometry-Guided Runge–Kutta Integration for Diffusion Transformer AccelerationChaoqun Sun, Zongjing Fu, Powei Chang, Jinpeng Zhang 等CVPR 2026
- PHAC: Promptable Human Amodal CompletionSeung Young Noh, Ju Yong ChangCVPR 2026
它引用的顶会 Paper35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- StreamDiffusion: A Pipeline-Level Solution for Real-Time Interactive GenerationAkio Kodaira, Chenfeng Xu, Toshiki Hazama, Takanori Yoshimoto 等ICCV 2025 · 被引用 10 次
- Generating compositional scenes via Text-to-image RGBA Instance GenerationAlessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang 等NeurIPS 2024 · 被引用 13 次
- Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic PromptsXinhua Cheng, Tianyu Yang, Jianan Wang, Yu Li 等ICLR 2024 · 被引用 58 次
- DialogDraw: Image Generation and Editing System Based on Multi-Turn DialogueShichao Ma, Xinfeng Zhang, Zeng Zhao, Bai Liu 等AAAI 2025 · 被引用 3 次
- Diffusion Texture PaintingAnita Hu, Nishkrit Desai, Hassan Abu Alhaija, Seung Wook Kim 等SIGGRAPH 2024 · 被引用 16 次
