Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing
Shichao Ma, Yunhe Guo, Jiahao Su, Qihe Huang, Zhengyang Zhou, Yang Wang
摘要
Text-to-image generation tasks have driven remarkable advances in diverse media applications, yet most focus on single-turn scenarios and struggle with iterative, multi-turn creative tasks. Recent dialogue-based systems attempt to bridge this gap, but their single-agent, sequential paradigm often causes intention drift and incoherent edits. To address these limitations, we present Talk2Image, a novel multiagent system for interactive image generation and editing in multi-turn dialogue scenarios. Our approach integrates three key components: intention parsing from dialogue history, task decomposition and collaborative execution across specialized agents, and feedback-driven refinement based on a multiview evaluation mechanism. Talk2Image enables step-bystep alignment with user intention and consistent image editing. Experiments demonstrate that Talk2Image outperforms existing baselines in controllability, coherence, and user satisfaction across iterative image generation and editing tasks. A black leather wallet lies on a marble countertop, and a credit card is partially sticking out of it. Remove the credit card from the wallet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ReCreate: Reasoning and Creating Domain Agents Driven by ExperienceZhezheng Hao, Hong Wang, Jian Luo, Jianqing Zhang 等ACL 2026 · 被引用 16 次
- SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM SystemsYuzhe Zhang, Feiran Liu, Yi Shan, Xinyi Huang 等ACL 2026 · 被引用 5 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Twin Co-Adaptive Dialogue for Progressive Image GenerationJianhui Wang, Yangfan He, Yan Zhong, Xinyuan Song 等ACM MM 2025 · 被引用 4 次
- CREA: A Collaborative Multi-Agent Framework for Creative Image Editing and GenerationKavana Venkatesh, Connor Dunlop, Pinar YanardagNeurIPS 2025 · 被引用 18 次
- Talk-to-Edit: Fine-Grained Facial Editing via DialogYuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy 等ICCV 2021 · 被引用 162 次
- LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object IntegrationYuyao Zhang, Jinghao Li, Yu-Wing TaiNeurIPS 2025 · 被引用 21 次
- Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in ScenesJing Tan, Zhaoyang Zhang, Yantao Shen, Jiarui Cai 等CVPR 2026 · 被引用 3 次
