Lune

ICLR2026顶会

Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models

Tianren Ma, Mu Zhang, Yibing Wang, Qixiang Ye

2026年份
10被引次数
3顶会引用

摘要

Optimizing discrete diffusion model (DDM) with rewards remains a challenge-the non-autoregressive paradigm makes importance sampling intractable and rollout complex, puzzling reinforcement learning methods such as Group Relative Policy Optimization (GRPO). In this study, we introduce MaskGRPO, the first viable approach to enable scalable multimodal reinforcement learning in discrete diffusion with effective importance sampling and modality-specific adaptations. To this end, we first clarify the theoretical foundation for DDMs, which facilitates building an importance estimator that captures valuable token fluctuation for gradient updates. We then delicately tailored the rollout method for visual sequences, which yields diverse completions and reliable optimization gradients. Upon math reasoning, coding, and visual generation benchmarks, MaskGRPO brings more stable and efficient updates, leading to stronger reasoning performance and better generation quality. This study establishes MaskGRPO as a systematic policy optimization approach and the first practical way for discretized visual diffusion. Our code is available at https://github.com/martian422/MaskGRPO . 𝜋𝜋 𝜃𝜃 𝑜𝑜𝑜𝑜𝑜𝑜 o 𝑖𝑖 o 𝑗𝑗 o 𝑘𝑘 � 𝜌𝜌 𝑗𝑗 𝑡𝑡 = exp ℓ 𝜋𝜋𝜃𝜃 𝑜𝑜 𝑗𝑗 𝑡𝑡 , 𝑜𝑜 𝑗𝑗 -ℓ 𝜋𝜋𝜃𝜃 𝑜𝑜𝑜𝑜𝑜𝑜 𝑜𝑜 𝑗𝑗 𝑡𝑡 , 𝑜𝑜 𝑗𝑗 𝜋𝜋 𝜃𝜃 o 𝑖𝑖 o 𝑗𝑗 o 𝑘𝑘 Semi-AR sampler Emerge sampler AR-like re-masked 𝑜𝑜 𝑗𝑗 𝑡𝑡 Random re-masked 𝑜𝑜 𝑗𝑗 𝑡𝑡

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper21

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖