Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective
Jingyang Ou, Jiaqi Han, Minkai Xu, Shaoxuan Xu, Jianwen Xie, Stefano Ermon, Yi Wu, Chongxuan Li
Abstract
Reinforcement Learning (RL) has proven highly effective for autoregressive language models, but adapting these methods to diffusion large language models (dLLMs) presents fundamental challenges. The core difficulty lies in likelihood approximation: while autoregressive models naturally provide token-level conditional probabilities essential for token-level RL objectives (e.g., GRPO), dLLMs generate sequences through iterative non-autoregressive denoising steps that lack this factorization. To address this fundamental mismatch, we propose ELBObased Sequence-level Policy Optimization (ESPO), a principled RL framework that treats entire sequence generation as a single action and uses the ELBO as a tractable sequence-level likelihood proxy. Our method incorporates per-token normalization of importance ratios and robust KL-divergence estimation to ensure stable large-scale training. Extensive experiments on mathematical reasoning, coding, and planning tasks demonstrate that ESPO significantly outperforms token-level baselines, achieving dramatic improvements of 20-40 points on the Countdown task, while maintaining consistent gains on math and coding benchmarks. Our approach establishes sequence-level optimization as a principled and empirically effective paradigm for RL in dLLMs. Our code is available at https://github.com/ML-GSAI/ESPO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94a10cdf-ef4a-4cce-8604-a12c761aae81Cited by top-tier papers5
- d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language ModelsLeyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu et al.ACL 2026 · 14 citations
- LightningRL: Breaking the Accuracy–Parallelism Trade-off of Block-wise dLLMs via Reinforcement LearningYanzhe Hu, Yijie Jin, Pengfei Liu, Kai Yu et al.ICML 2026 · 5 citations
- Lavida-R1: Advancing Reasoning for Unified Multimodal Diffusion Language ModelsShufan Li, Yuchen Zhu, Kangning Liu, Zhe Lin et al.ICML 2026 · 4 citations
- The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language ModelsZanlin Ni, Shenzhi Wang, Yang Yue, Tianyu Yu et al.ICML 2026 · 4 citations
- Stabilizing Reinforcement Learning for Diffusion Language ModelsJianyuan Zhong, Wang Kaibo, Ding Ding, Zijin Feng et al.ICML 2026 · 3 citations
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
Related papers
- Improving Reasoning for Diffusion Language Models via Group Diffusion Policy OptimizationKevin Rojas, Jiahe Lin, Kashif Rasul, Anderson Schneider et al.ICLR 2026 · 34 citations
- Simple Policy Gradients for Reasoning with Diffusion Language ModelsAnthony ZhanICML 2026 · 4 citations
- SPG: Sandwiched Policy Gradient for Masked Diffusion Language ModelsChenyu Wang, Paria Rashidinejad, DiJia Andy Su, Song Jiang et al.ICLR 2026 · 45 citations
- TA-GRPO-d: Trajectory-Aware GRPO for Optimizing Denoising Trajectories in Diffusion LLMsGyunyeop Kim, Sangwoo KangACL 2026
- Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language ModelsNianyi Lin, Jiajie Zhang, Lei Hou, Juanzi LiACL 2026 · 8 citations
