Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
Shufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin, Zijun Wei, Aditya Grover, Jason Kuen
Abstract
We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks and low-resolution image generation, Lavida-O presents a single framework that enables image-level understanding, object grounding, image editing, and high-resolution (1024px) text-to-image synthesis. Lavida-O incorporates a novel Elastic Mixture-of-Transformers (Elastic-MoT) architecture that couples a lightweight generation branch with a larger understanding branch, supported by token compression, universal text conditioning and stratified sampling for efficient and high-quality generation. Lavida-O further incorporates planning and iterative self-reflection in image generation and editing tasks, seamlessly boosting generation quality with its understanding capabilities. Lavida-O achieves state-of-the-art performance on a wide range of benchmarks including RefCOCO object grounding, GenEval text-to-image generation, and ImgEdit image editing, outperforming existing autoregressive models and continuous diffusion models such as Qwen2.5-VL and FluxKontext-dev, while offering considerable speedup at inference. These advances establish Lavida-O as a new paradigm for scalable multimodal reasoning and generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c720b260-05cd-4fa4-b41b-c67445039182Cited by top-tier papers5
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and GenerationXiang Wang, Zhifei Zhang, He Zhang, Zhe Lin et al.CVPR 2026 · 12 citations
- Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language ModelsShufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin et al.CVPR 2026 · 6 citations
- LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language ModelsChenglin Wang, Yucheng Zhou, Shuang Chen, Tao Wang et al.ACL 2026 · 1 citation
- Efficient Test-Time Scaling via Hierarchical Search and Self-Verification for Discrete Diffusion Language ModelsJinbin Bai, Yixuan Li, Yuchen Zhu, Yi Xin et al.ICML 2026
- Adversarial Reinforcement Learning for Robust Diffusion Large Language Model UnlearningZhiwei Zhang, Yudi Lin, Linlin Wu, Fali Wang et al.ICML 2026
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- LaViDa: A Large Diffusion Language Model for Multimodal UnderstandingShufan Li, Konstantinos Kallidromitis, Hritik Bansal, Akash Gokul et al.NeurIPS 2025 · 89 citations
- Show-o: One Single Transformer to Unify Multimodal Understanding and GenerationJinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang et al.ICLR 2025
- Lavida-R1: Advancing Reasoning for Unified Multimodal Diffusion Language ModelsShufan Li, Yuchen Zhu, Kangning Liu, Zhe Lin et al.ICML 2026 · 4 citations
- Beyond Text-to-Image: Liberating Generation with a Unified Discrete Diffusion ModelQingyu Shi, Jinbin Bai, Zhuoran Zhao, Wenhao Chai et al.ICLR 2026 · 40 citations
- MMaDA: Multimodal Large Diffusion Language ModelsLing Yang, Ye Tian, Bowen Li, Xinchen Zhang et al.NeurIPS 2025 · 255 citations
