DialogDraw: Image Generation and Editing System Based on Multi-Turn Dialogue
Shichao Ma, Xinfeng Zhang, Zeng Zhao, Bai Liu, Changjie Fan, Zhipeng Hu
摘要
In recent years, diffusion modeling has shown great potential for image generation and editing. Beyond single-model approaches, various drawing workflows now exist to handle diverse drawing tasks. However, few solutions effectively identify user intentions through dialogue and progressively complete drawings. We introduce DialogDraw, which facilitates image generation and editing through continuous dialogue interaction. DialogDraw enables users to create and refine drawings using natural language and integrates with numerous open-source drawing workflows and models. The system accurately recognizes intentions and extracts user inputs via parameterization, adapts to various drawing function parameters, and provides an intuitive interaction mode. It effectively executes user instructions, supports dozens of image generation and editing methods, and offers robust scalability. Moreover, we employ SFT and RLHF to iterate the Intention Recognition and Parameter Extraction Model (IRPEM). To evaluate DialogDraw's functionality, we propose DrawnConvos, a dataset rich in drawing functions and command dialogue data collected from the open-source community. Our evaluation demonstrates that DialogDraw excels in command compliance, identifying and adapting to user drawing intentions, thereby proving the effectiveness of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Charts Are Not Images: On the Challenges of Scientific Chart EditingShawn Li, Ryan Rossi, Sungchul Kim, Sunav Choudhary 等ICLR 2026 · 被引用 11 次
- Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and EditingShichao Ma, Yunhe Guo, Jiahao Su, Qihe Huang 等AAAI 2026 · 被引用 9 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion ModelsJaerin Lee, Daniel Sungho Jung, Kanggeon Lee, Kyoung Mu LeeCVPR 2025
- MultiFusion: Fusing Pre-Trained Models for Multi-Lingual, Multi-Modal Image GenerationMarco Bellagente, Manuel Brack, Hannah Teufel, Felix Friedrich 等NeurIPS 2023 · 被引用 31 次
- Pattern Analogies: Learning to Perform Programmatic Image Edits by AnalogyAditya Ganeshan, Thibault Groueix, Paul Guerrero, Radomír Mech 等CVPR 2025
- Make Me Happier: Evoking Emotions through Image Diffusion ModelsQing Lin, Jingfeng Zhang, Yew-Soon Ong, Mengmi ZhangICCV 2025 · 被引用 4 次
- CoProSketch: Controllable and Progressive Sketch Generation with Diffusion ModelRuohao Zhan, Yijin Li, Yisheng He, Shuo Chen 等ACM MM 2025 · 被引用 2 次
