DreamOmni: Unified Image Generation and Editing
Bin Xia, Yuechen Zhang, Jingyao Li, Chengyao Wang, Yitong Wang, Xinglong Wu, Bei Yu, Jiaya Jia
Abstract
projectpagepCurrently, the success of large language models (LLMs) illustrates that a unified multitasking approach can significantly enhance model usability, stream-line deployment, and foster synergistic benefits across different tasks. However, in computer vision, while text-to-image (T2I) models have significantly improved generation quality through scaling up, their framework design did not initially consider how to unify with downstream tasks, such as various types of editing. To address this, we introduce DreamOmni, a unified model for image generation and editing. We begin by analyzing existing frameworks and the requirements of downstream tasks, proposing a unified framework that integrates both T2I models and various editing tasks. Furthermore, another key challenge is the efficient creation of high-quality editing data, particularly for instruction-based and drag-based editing. To this end, we develop a synthetic data pipeline using sticker-like elements to synthesize accurate, high-quality datasets efficiently, which enables editing data scaling up for unified model training. For training, DreamOmni jointly trains T2I generation and downstream tasks. T2I training enhances the model’s understanding of specific concepts and improves generation quality, while editing training helps the model grasp the nuances of the editing task. This collaboration significantly boosts editing performance. Extensive experiments confirm the effectiveness of DreamOmni. The code and model will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 930e56d6-3b5b-48ff-822e-1efa358d88dcCited by top-tier papers8
- iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image GenerationZhoujie Fu, Xianfang Zeng, Jinghong Lan, Xinyao Liao et al.CVPR 2026 · 7 citations
- The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive AlignmentZiheng Ouyang, Yiren Song, Yaoli Liu, Shihao Zhu et al.CVPR 2026 · 6 citations
- VDOT: Efficient Unified Video Creation via Optimal Transport DistillationYutong Wang, Haiyu Zhang, Tianfan Xue, Yu Qiao et al.CVPR 2026 · 5 citations
- Mixture-of-Scores: Robust Image-Text Data Valuation via Three Lines of CodeSitong Wu, Haoru Tan, Yukang Chen, Shaofeng Zhang et al.ICCV 2025 · 4 citations
- Omni2: Unifying Omnidirectional Image Generation and Editing in an Omni ModelLiu Yang, Huiyu Duan, Yucheng Zhu, Xiaohong Liu et al.ACM MM 2025 · 2 citations
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- DreamOmni2: Multimodal Instruction-based Generation and EditingBin Xia, Bohao Peng, Yuechen Zhang, Junjia Huang et al.CVPR 2026
- OmniGen: Unified Image GenerationShitao Xiao, Yueze Wang, Junjie Zhou, Huaying Yuan et al.CVPR 2025
- Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation TasksRuibin Li, Tao Yang, Yangming Shi, Weiguo Feng et al.ICLR 2026 · 4 citations
- Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and EditingZeyue Tian, Binxin Yang, Zhaoyang Liu, Jiexuan Zhang et al.SIGGRAPH 2026
- MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and EditingXueyun Tian, Wei Li, Bingbing Xu, Yige Yuan et al.ACM MM 2025 · 4 citations
