Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing
Shichao Ma, Yunhe Guo, Jiahao Su, Qihe Huang, Zhengyang Zhou, Yang Wang
Abstract
Text-to-image generation tasks have driven remarkable advances in diverse media applications, yet most focus on single-turn scenarios and struggle with iterative, multi-turn creative tasks. Recent dialogue-based systems attempt to bridge this gap, but their single-agent, sequential paradigm often causes intention drift and incoherent edits. To address these limitations, we present Talk2Image, a novel multiagent system for interactive image generation and editing in multi-turn dialogue scenarios. Our approach integrates three key components: intention parsing from dialogue history, task decomposition and collaborative execution across specialized agents, and feedback-driven refinement based on a multiview evaluation mechanism. Talk2Image enables step-bystep alignment with user intention and consistent image editing. Experiments demonstrate that Talk2Image outperforms existing baselines in controllability, coherence, and user satisfaction across iterative image generation and editing tasks. A black leather wallet lies on a marble countertop, and a credit card is partially sticking out of it. Remove the credit card from the wallet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- ReCreate: Reasoning and Creating Domain Agents Driven by ExperienceZhezheng Hao, Hong Wang, Jian Luo, Jianqing Zhang et al.ACL 2026 · 16 citations
- SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM SystemsYuzhe Zhang, Feiran Liu, Yi Shan, Xinyi Huang et al.ACL 2026 · 5 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Twin Co-Adaptive Dialogue for Progressive Image GenerationJianhui Wang, Yangfan He, Yan Zhong, Xinyuan Song et al.ACM MM 2025 · 4 citations
- CREA: A Collaborative Multi-Agent Framework for Creative Image Editing and GenerationKavana Venkatesh, Connor Dunlop, Pinar YanardagNeurIPS 2025 · 18 citations
- Talk-to-Edit: Fine-Grained Facial Editing via DialogYuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy et al.ICCV 2021 · 162 citations
- LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object IntegrationYuyao Zhang, Jinghao Li, Yu-Wing TaiNeurIPS 2025 · 21 citations
- Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in ScenesJing Tan, Zhaoyang Zhang, Yantao Shen, Jiarui Cai et al.CVPR 2026 · 3 citations
