Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
Peng Sun, Jun XIE, Tao Lin
Abstract
Unified Multimodal Models (UMMs) are often constrained by the pre-training of their , which typically relies on inefficient paradigms and scarce, high-quality text-image paired data. In this paper, we systematically analyze pre-training recipes for and identify these two issues as the major bottlenecks. To address them, we propose , a data-efficient two-stage training framework. The first stage pre-trains the visual generative component using abundant unlabeled image-only data, thereby removing the dependency on paired data . The second stage fine-tunes the model using a mixture of unlabeled images and a small curated set of text-image pairs, leading to improved instruction alignment and generative quality. Extensive experiments show that IOMM not only improves training efficiency but also achieves state-of-the-art performance. For example, our IOMM-B (3.6B) model was trained from scratch using only H800 GPU hours (with the vast majority, hours, dedicated to the efficient ). It achieves on GenEval and on WISE surpassing strong baselines such as BAGEL-7B (0.82 & 0.55) and BLIP3-o-4B (0.84 & 0.50). Code will be released publicly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28e52a62-bd7b-4f74-a690-cd66d1291aa6Builds on26
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- ILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhanceChunwei Wang, Guansong Lu, Junwei Yang, Runhui Huang et al.ICCV 2025 · 5 citations
- UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningWei Li, Can Gao, Guocheng Niu, Xinyan Xiao et al.ACL 2021
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal ModelsJitai Hao, Hao Liu, Xinyan Xiao, Qiang Huang et al.ICLR 2026 · 18 citations
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and GenerationRui Tian, Mingfei Gao, Mingze Xu, Jiaming Hu et al.NeurIPS 2025 · 35 citations
- Reconstruction Alignment Improves Unified Multimodal ModelsJi Xie, Trevor Darrell, Luke Zettlemoyer, XuDong WangICLR 2026 · 52 citations
