DialogDraw: Image Generation and Editing System Based on Multi-Turn Dialogue
Shichao Ma, Xinfeng Zhang, Zeng Zhao, Bai Liu, Changjie Fan, Zhipeng Hu
Abstract
In recent years, diffusion modeling has shown great potential for image generation and editing. Beyond single-model approaches, various drawing workflows now exist to handle diverse drawing tasks. However, few solutions effectively identify user intentions through dialogue and progressively complete drawings. We introduce DialogDraw, which facilitates image generation and editing through continuous dialogue interaction. DialogDraw enables users to create and refine drawings using natural language and integrates with numerous open-source drawing workflows and models. The system accurately recognizes intentions and extracts user inputs via parameterization, adapts to various drawing function parameters, and provides an intuitive interaction mode. It effectively executes user instructions, supports dozens of image generation and editing methods, and offers robust scalability. Moreover, we employ SFT and RLHF to iterate the Intention Recognition and Parameter Extraction Model (IRPEM). To evaluate DialogDraw's functionality, we propose DrawnConvos, a dataset rich in drawing functions and command dialogue data collected from the open-source community. Our evaluation demonstrates that DialogDraw excels in command compliance, identifying and adapting to user drawing intentions, thereby proving the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0811222a-a5d3-4695-bfdd-21a8bc453b0fCited by top-tier papers2
- Charts Are Not Images: On the Challenges of Scientific Chart EditingShawn Li, Ryan Rossi, Sungchul Kim, Sunav Choudhary et al.ICLR 2026 · 11 citations
- Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and EditingShichao Ma, Yunhe Guo, Jiahao Su, Qihe Huang et al.AAAI 2026 · 9 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion ModelsJaerin Lee, Daniel Sungho Jung, Kanggeon Lee, Kyoung Mu LeeCVPR 2025
- MultiFusion: Fusing Pre-Trained Models for Multi-Lingual, Multi-Modal Image GenerationMarco Bellagente, Manuel Brack, Hannah Teufel, Felix Friedrich et al.NeurIPS 2023 · 31 citations
- Pattern Analogies: Learning to Perform Programmatic Image Edits by AnalogyAditya Ganeshan, Thibault Groueix, Paul Guerrero, Radomír Mech et al.CVPR 2025
- Make Me Happier: Evoking Emotions through Image Diffusion ModelsQing Lin, Jingfeng Zhang, Yew-Soon Ong, Mengmi ZhangICCV 2025 · 4 citations
- CoProSketch: Controllable and Progressive Sketch Generation with Diffusion ModelRuohao Zhan, Yijin Li, Yisheng He, Shuo Chen et al.ACM MM 2025 · 2 citations
