wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
Xiaohang Tang, Rares Dolga, Sangwoong Yoon, Ilija Bogunovic
摘要
Improving the reasoning capabilities of diffusion-based large language models (dLLMs) through reinforcement learning (RL) remains an open problem. The intractability of dLLMs likelihood function necessitates approximating the current, old, and reference policy likelihoods at each policy optimization step. This reliance introduces additional computational overhead, and can lead to large variance and estimation error in RL objective -- particularly in computing the policy ratio for importance sampling. To mitigate these issues, we introduce wd1, a novel ratio-free policy optimization approach that reformulates the RL objective as a weighted log-likelihood, requiring only a single approximation for the current parametrized policy likelihood. We formally show that our proposed method can be interpreted as energy-guided discrete diffusion training combined with negative sample unlearning, thereby confirming its theoretical soundness. In experiments on LLaDA-8B model, wd1 outperforms diffusion-based GRPO (d1) while requiring lower computational cost, achieving up to a +59% improvement in accuracy. Furthermore, we extend wd1 to denoising-stepwise weighted policy optimization (wd1++), achieving state-of-the-art math performance of 44.2% on MATH500 and 84.5% on GSM8K with only 20 RL training steps.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy OptimizationYuchen Zhu, Wei Guo, Jaemoo Choi, Petr Molodyk 等ICML 2026 · 被引用 13 次
- Consolidating Reinforcement Learning for Multimodal Discrete Diffusion ModelsTianren Ma, Mu Zhang, Yibing Wang, Qixiang YeICLR 2026 · 被引用 10 次
- Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language ModelsNianyi Lin, Jiajie Zhang, Lei Hou, Juanzi LiACL 2026 · 被引用 8 次
- Breaking the Factorization Barrier in Diffusion Language ModelsIan Li, Zilei Shao, Benjie Wang, Rose Yu 等ICML 2026 · 被引用 5 次
- Lavida-R1: Advancing Reasoning for Unified Multimodal Diffusion Language ModelsShufan Li, Yuchen Zhu, Kangning Liu, Zhe Lin 等ICML 2026 · 被引用 4 次
它引用的顶会 Paper30
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
相关 Paper
- Stabilizing Reinforcement Learning for Diffusion Language ModelsJianyuan Zhong, Wang Kaibo, Ding Ding, Zijin Feng 等ICML 2026 · 被引用 3 次
- Principled RL for Diffusion LLMs Emerges from a Sequence-Level PerspectiveJingyang Ou, Jiaqi Han, Minkai Xu, Shaoxuan Xu 等ICLR 2026 · 被引用 33 次
- d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement LearningSiyan Zhao, Devaansh Gupta, Qinqing Zheng, Aditya GroverNeurIPS 2025 · 被引用 191 次
- Simple Policy Gradients for Reasoning with Diffusion Language ModelsAnthony ZhanICML 2026 · 被引用 4 次
- The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language ModelsZanlin Ni, Shenzhi Wang, Yang Yue, Tianyu Yu 等ICML 2026 · 被引用 4 次
