EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models
Rui Zhao, Hangjie Yuan, Yujie Wei, Shiwei Zhang, Yuchao Gu, Lingmin Ran, Xiang Wang, Jay Zhangjie Wu, David Junhao Zhang, Yingya Zhang, Mike Zheng Shou
Abstract
Recent advancements in generation models have showcased remarkable capabilities in generating fantastic content. However, most of them are trained on proprietary high-quality data, and some models withhold their parameters and only provide accessible application programming interfaces (APIs), limiting their benefits for downstream tasks. To explore the feasibility of training a text-to-image generation model comparable to advanced models using publicly available resources, we introduce EvolveDirector. This framework interacts with advanced models through their public APIs to obtain text-image data pairs to train a base model. Our experiments with extensive data indicate that the model trained on generated data of the advanced model can approximate its generation capability. However, it requires large-scale samples of 10 million or more. This incurs significant expenses in time, computational resources, and especially the costs associated with calling fee-based APIs. To address this problem, we leverage pre-trained large vision-language models (VLMs) to guide the evolution of the base model. VLM continuously evaluates the base model during training and dynamically updates and refines the training dataset by the discrimination, expansion, deletion, and mutation operations. Experimental results show that this paradigm significantly reduces the required data volume. Furthermore, when approaching multiple advanced models, EvolveDirector can select the best samples generated by them to learn powerful and balanced abilities. The final trained model Edgen is demonstrated to outperform these advanced models. The code and model weights are available at https://github.com/showlab/EvolveDirector.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b8a991b-1e78-4fda-bd76-ee31d83d90faCited by top-tier papers5
- Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation TrainingZiqi Gao, Weikai Huang, Jieyu Zhang, Aniruddha Kembhavi et al.ICLR 2026 · 1 citation
- Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-LearningTrinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen et al.KDD 2026 · 1 citation
- RAVEN: Realtime Accessibility in Virtual ENvironments for Blind and Low-Vision PeopleXinyun Cao, Kexin Phyllis Ju, Chenglin Li, Venkatesh Potluri et al.CHI 2026 · 1 citation
- P-Flow: Prompting Visual Effects GenerationRui Zhao, Mike Zheng ShouCVPR 2026
- DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal CyclesRui Zhao, Weijia Mao, Mike Zheng ShouCVPR 2025
Builds on34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Expedited Training of Visual Conditioned Language Generation via Redundancy ReductionYiren Jian, Tingkai Liu, Yunzhe Tao, Chunhui Zhang et al.ACL 2024 · 5 citations
- Visual Programming for Step-by-Step Text-to-Image Generation and EvaluationJaemin Cho, Abhay Zala, Mohit BansalNeurIPS 2023 · 62 citations
- Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data GenerationYue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta et al.ACL 2025
- ProAPO: Progressively Automatic Prompt Optimization for Visual ClassificationXiangyan Qu, Gaopeng Gou, Jiamin Zhuang, Jing Yu et al.CVPR 2025
- Learning an Image Editing Model without Image Editing PairsNupur Kumari, Sheng-Yu Wang, Nanxuan Zhao, Yotam Nitzan et al.ICLR 2026 · 14 citations
