GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing
Zhenyu Wang, Aoxue Li, Zhenguo Li, Xihui Liu
Abstract
Despite the success achieved by existing image generation and editing methods, current models still struggle with complex problems including intricate text prompts, and the absence of verification and self-correction mechanisms makes the generated images unreliable. Meanwhile, a single model tends to specialize in particular tasks and possess the corresponding capabilities, making it inadequate for fulfilling all user requirements. We propose GenArtist, a unified image generation and editing system, coordinated by a multimodal large language model (MLLM) agent. We integrate a comprehensive range of existing models into the tool library and utilize the agent for tool selection and execution. For a complex problem, the MLLM agent decomposes it into simpler sub-problems and constructs a tree structure to systematically plan the procedure of generation, editing, and self-correction with step-by-step verification. By automatically generating missing position-related inputs and incorporating position information, the appropriate tool can be effectively employed to address each sub-problem. Experiments demonstrate that GenArtist can perform various generation and editing tasks, achieving state-of-the-art performance and surpassing existing models such as SDXL and DALL-E 3, as can be seen in Fig. 1. Project page is https://zhenyuw16.github.io/GenArtist_page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efe449b2-62e9-4754-8cdf-5a4473923bb6Cited by top-tier papers50
- ImageRAG: Dynamic Image Retrieval for Reference-Guided Image GenerationRotem Shalev-Arkushin, Rinon Gal, Amit Bermano, Ohad FriedICLR 2026 · 25 citations
- LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object IntegrationYuyao Zhang, Jinghao Li, Yu-Wing TaiNeurIPS 2025 · 21 citations
- RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement LearningMingrui Wu, Lu Wang, Pu Zhao, Fangkai Yang et al.ICLR 2026 · 19 citations
- ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive FeedbackLitao Guo, Xinli Xu, Luozhou Wang, Jiantao Lin et al.NeurIPS 2025 · 18 citations
- CREA: A Collaborative Multi-Agent Framework for Creative Image Editing and GenerationKavana Venkatesh, Connor Dunlop, Pinar YanardagNeurIPS 2025 · 18 citations
Builds on38
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- GENMAC: Compositional Text-to-Video Generation with Multi-Agent CollaborationKaiyi Huang, Yukun Huang, Xuefei Ning, Zinan Lin et al.AAAI 2026 · 1 citation
- Hybrid Agents for Image RestorationBingchen Li, Xin Li, Yiting Lu, Zhibo ChenCVPR 2026 · 17 citations
- Muses: 3D-Controllable Image Generation via Multi-Modal Agent CollaborationYanbo Ding, Shaobin Zhuang, Kunchang Li, Zhengrong Yue et al.AAAI 2025 · 8 citations
- GenAssist: Making Image Generation AccessibleMina Huh, Yi-Hao Peng, Amy PavelUIST 2023 · 58 citations
- MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching SkillsNiladri Shekhar Dutt, Duygu Ceylan, Niloy J. MitraSIGGRAPH 2025 · 2 citations
