TA-GRPO-d: Trajectory-Aware GRPO for Optimizing Denoising Trajectories in Diffusion LLMs
Gyunyeop Kim, Sangwoo Kang
Abstract
Diffusion large language models (dLLMs) generate text by repeatedly unmasking a partially noised sequence in parallel, promising lower latency than autoregressive decoding. However, most discrete dLLMs still rely on fixed denoising schedules, which are non-adaptive to input difficulty and cannot learn efficient unmasking orders. This paper introduces a reinforcement learning (RL) framework that transforms dLLM decoding into a trajectory-aware, learnable policy. We propose a confidence-gated denoising strategy that dynamically decides which tokens to unmask and how many to un-mask per step, enabling adaptive exploration of denoising trajectories. Building on Group Relative Policy Optimization, we reformulate it into a trajectory-aware variant, TA-GRPO-d , which combines a trajectory-level signal—captured as the z-score of the AUC over intermediate rewards—with a token-level unmasking-time weight. This design allows the model to learn not only the final output quality but also the efficiency of the decoding path itself. Experiments on MATH-500, Countdown, Sudoku, and code benchmarks (HumanEval, MBPP) show that TA-GRPO-d maintains or improves accuracy while reducing average denoising steps by up to half, achieving both faster inference and lower computational cost. Our approach provides an RL framework for optimizing dLLM decoding policies toward adaptive, efficient reasoning. Code is available at our GitHub 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2c366c4-86b6-45f7-baa7-031011f78ca7Builds on9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
Related papers
- The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language ModelsZanlin Ni, Shenzhi Wang, Yang Yue, Tianyu Yu et al.ICML 2026 · 4 citations
- Learning Unmasking Policies for Diffusion Language ModelsMetod Jazbec, Theo X. Olausson, Louis Béthune, Pierre Ablin et al.ICML 2026 · 24 citations
- Simple Policy Gradients for Reasoning with Diffusion Language ModelsAnthony ZhanICML 2026 · 4 citations
- Principled RL for Diffusion LLMs Emerges from a Sequence-Level PerspectiveJingyang Ou, Jiaqi Han, Minkai Xu, Shaoxuan Xu et al.ICLR 2026 · 33 citations
- Improving Reasoning for Diffusion Language Models via Group Diffusion Policy OptimizationKevin Rojas, Jiahe Lin, Kashif Rasul, Anderson Schneider et al.ICLR 2026 · 34 citations
