MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
Ye Tian, Ling Yang, Jiongfan Yang, Anran Wang, Yu Tian, Jiani Zheng, Haochen Wang, Zhiyang Teng, Zhuochen Wang, Yinjie Wang, Yunhai Tong, Mengdi Wang, Xiangtai Li
摘要
While thinking-aware generation aims to improve performance on complex tasks, we identify a critical failure mode where existing sequential, autoregressive approaches can paradoxically degrade performance due to error propagation. To systematically analyze this issue, we propose ParaBench, a new benchmark designed to evaluate both text and image output modalities. Our analysis using ParaBench reveals that this performance degradation is strongly correlated with poor alignment between the generated reasoning and the final image. To resolve this, we propose a parallel multimodal diffusion framework, MMaDA-Parallel, that enables continuous, bidirectional interaction between text and images throughout the entire denoising trajectory. MMaDA-Parallel is trained with supervised finetuning and then further optimized by Parallel Reinforcement Learning (ParaRL), a novel strategy that applies semantic rewards along the trajectory to enforce cross-modal consistency. Experiments validate that our model significantly improves cross-modal alignment and semantic consistency, achieving a 6.9% improvement in Output Alignment on ParaBench compared to the state-of-the-art model, Bagel, establishing a more robust paradigm for thinking-aware image synthesis. Our code is open-sourced at https://github.com/tyfeld/MMaDA-Parallel
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual GenerationJunjie Wang, 星华 娄, Xiangtai Li, Ye Tian 等ICML 2026
- Adversarial Reinforcement Learning for Robust Diffusion Large Language Model UnlearningZhiwei Zhang, Yudi Lin, Linlin Wu, Fali Wang 等ICML 2026
它引用的顶会 Paper31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and PerceptionXinyang Song, Libin Wang, Weining Wang, Shaozhen Liu 等AAAI 2026
- Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM EncodersSiqi Kou, Jiachun Jin, Zetong Zhou, YE MA 等ICML 2026 · 被引用 13 次
- MMaDA: Multimodal Large Diffusion Language ModelsLing Yang, Ye Tian, Bowen Li, Xinchen Zhang 等NeurIPS 2025 · 被引用 255 次
- DiffThinker: Towards Generative Multimodal Reasoning with Diffusion ModelsZefeng He, Xiaoye Qu, Yafu Li, Tong Zhu 等ICML 2026 · 被引用 13 次
- Benchmarking and Improving Fine-Grained Text-to-Image Alignment via Paired Reinforcement LearningKaihang Pan, Wendong Bu, Yuruo Wu, Kai Shen 等ICML 2026
