Lune

CVPR2025Top-tier venue

Dual Diffusion for Unified Image Generation and Understanding

Zijie Li, Henry Li, Yichun Shi, Amir Barati Farimani, Yuval Kluger, Linjie Yang, Peng Wang

2025Year
41Top-tier citations

Abstract

Diffusion models have gained tremendous success in textto-image generation, yet still struggle with visual understanding tasks, an area dominated by autoregressive visionlanguage models. We propose a large-scale and fully endto-end diffusion model for multi-modal understanding and generation that significantly improves on existing diffusionbased multimodal models, and is the first of its kind to support the full suite of vision-language modeling capabilities. Inspired by the multimodal diffusion transformer (MM-DiT) and recent advances in discrete diffusion language modeling, we leverage a cross-modal maximum likelihood estimation framework that simultaneously trains the conditional likelihoods of both images and text jointly under a single loss function, which is back-propagated through both branches of the diffusion transformer. The resulting model is highly flexible and capable of a wide range of tasks including image generation, captioning, and visual question answering. Our model attained competitive performance compared to recent unified image understanding and generation models, demonstrating the potential of multimodal diffusion modeling as a promising alternative to autoregressive next-token prediction models.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3d69314b-4cce-4953-89ef-64f63a7ea7aa

Cited by top-tier papers41

Ask how each one uses it

Builds on42

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines