Lavida-R1: Advancing Reasoning for Unified Multimodal Diffusion Language Models
Shufan Li, Yuchen Zhu, Kangning Liu, Zhe Lin, Yongxin Chen, Molei Tao, Aditya Grover, Jiuxiang Gu, Jason Kuen
Abstract
Diffusion language models (dLLMs) recently emerged as a promising alternative to auto-regressive LLMs. The latest works further extended it to multimodal understanding and generation tasks. In this work, we propose LaViDa-R1, a multimodal, general-purpose reasoning dLLM. Unlike existing works that build reasoning dLLMs through task-specific reinforcement learning, LaViDa-R1 incorporates diverse multimodal understanding and generation tasks in a unified manner. In particular, LaViDa-R1 is built with a novel unified post-training framework that seamlessly integrates supervised finetuning (SFT) and multi-task reinforcement learning (RL). It employs several novel training techniques, including answer-forcing, tree search, and complementary likelihood estimation, to enhance effectiveness and scalability. Extensive experiments demonstrate LaViDa-R1's strong performance on a wide range of multimodal tasks, including visual math reasoning, reason-intensive grounding, and image editing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fdaef03-492e-4637-bb01-40ef2b0f9098Cited by top-tier papers1
Ask how each one uses itBuilds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
Related papers
- LaViDa: A Large Diffusion Language Model for Multimodal UnderstandingShufan Li, Konstantinos Kallidromitis, Hritik Bansal, Akash Gokul et al.NeurIPS 2025 · 89 citations
- Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and GenerationShufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin et al.ICLR 2026 · 14 citations
- MMaDA: Multimodal Large Diffusion Language ModelsLing Yang, Ye Tian, Bowen Li, Xinchen Zhang et al.NeurIPS 2025 · 255 citations
- d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement LearningSiyan Zhao, Devaansh Gupta, Qinqing Zheng, Aditya GroverNeurIPS 2025 · 191 citations
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning ModelTianle Li, Jihai Zhang, Yongming Rao, Yu ChengNeurIPS 2025 · 17 citations
